Mitigating Catastrophic Forgetting In Reinforcement Learning.pdf

1802.07239.pdf
Preview of Mitigating Catastrophic Forgetting in Reinforcement Learning
🔗 Source: deepsense.ai
📊 Size: 2.49 MB
👤 Author: Christos Kaplanis, Murray Shanahan, Claudia Clopath
⬇️ Downloads: 204

Summary

Unlike humans, who can learn continuously over their lifetimes, artificial neural networks suffer from catastrophic forgetting, where new learning can abruptly erase previously acquired knowledge. Unlike neural networks, where parameters are typically modeled as scalar values, individual synapses in the brain comprise a complex network of interacting biochemical components that evolve at different timescales. By equipping tabular and deep reinforcement learning agents with a synaptic model that incorporates this biological complexity, catastrophic forgetting can be mitigated at multiple timescales. This model enables continual learning across sequential training of two simple tasks and can also be used to overcome within-task forgetting by reducing the need for an experience replay database.

Synaptic plasticity, the ability of connections between neurons to change their strength over time, is widely considered the physical basis of learning in the brain. Knowledge is thought to be distributed across neuronal networks, with individual synapses participating in the storage of several memories. Given this overlapping nature of memory storage, synapses need to be both labile in response to new experiences and stable enough to retain old memories, a paradox often referred to as the stability-plasticity dilemma.

Artificial neural networks also have a distributed memory but, unlike the brain, are prone to catastrophic forgetting when trained on a nonstationary data distribution. In reinforcement learning, where data is typically accumulated online as the agent interacts with the environment, the distribution of experiences is often nonstationary over the training of a single task, as well as across tasks. A typical way of addressing nonstationarity of data in deep RL is to store experiences in a replay database and use it to interleave old data and new data during training. However, this solution does not scale well computationally as the number of tasks grows and the old data might also become unavailable at some point.

One potential answer to how the brain achieves continual learning may arise from the experimental observations that synaptic plasticity occurs at a range of different timescales, including short-term plasticity, long-term plasticity, and synaptic consolidation. Intuitively, the slow components to plasticity could ensure that a synapse retains memory of a long history of its modifications, while the fast components render the synapse highly adaptable to the formation of new memories, perhaps providing a solution to the stability-plasticity dilemma.

In this paper, we explore whether a biologically plausible synaptic model, which abstractly models plasticity over a range of timescales, can be applied to mitigate catastrophic forgetting in a reinforcement learning context. Our work is intended as a proof of principle for how the incorporation of biological complexity to an agent's parameters can be useful in tackling the lifelong learning problem. By running experiments with both tabular and deep RL agents, we find that the model helps continual learning across two simple tasks as well as within a single task, by allaying the necessity of an experience replay database.

The Benna-Fusi model, which we make use of, assumes that a synaptic weight w at time t is determined by its history of modifications up until that time ∆w(t′), which are filtered by some kernel r(t −t′). The model abstracts away from the causes of the synaptic modifications ∆w and so is amenable for testing in different learning settings. The model consists of a finite chain of N communicating dynamic variables, where the dynamics of each variable uk in the chain are determined by interaction with its neighbours in the chain. The synaptic weight itself w is just read off from the value of u1, while the other variables are hidden and have the effect of regularising the value of the weight by the history of its modifications.

From a mechanical perspective, one can draw a comparison between the dynamics of the chain of variables and liquid flowing through a series of beakers with different base areas Ck connected by tubes of widths gk−1,k and gk,k+1. The value of a uk variable corresponds to the level of liquid in the beaker. Given a finite number of beakers per synapse, the best approximation to a power law decay is achieved by exponentially increasing the base areas of the beakers and exponentially decreasing the tube widths as you move down the chain. Beakers with wide bases and connected by smaller tubes will necessarily evolve at longer timescales. From a biological perspective, the dynamic variables can be likened to reversible biochemical processes that are related to plasticity and occur at a large range of timescales.

Description

Unlike humans, artificial neural networks suffer from catastrophic forgetting, where new learning erases previously acquired knowledge. We show that a synaptic model with biological complexity can mitigate this issue at multiple timescales. This enables continual learning across sequential training of tasks.

Technical Information

  • File Format: PDF
  • File Size: 2.49 MB
  • Pages: 14
  • Language: EN
  • Author: Christos Kaplanis, Murray Shanahan, Claudia Clopath
  • Total Downloads: 204
  • Last Updated: 1 week ago

Document Overview

This PDF document about Mitigating Catastrophic Forgetting in Reinforcement Learning provides comprehensive information and guidance. Whether you're a beginner or advanced user, this resource offers valuable insights into Mitigating Catastrophic Forgetting in Reinforcement Learning.

Related Topics

If you're interested in Mitigating Catastrophic Forgetting in Reinforcement Learning, you might also want to explore:

Download Mitigating Catastrophic Forgetting in Reinforcement Learning eBooks for free and learn more about Mitigating Catastrophic Forgetting in Reinforcement Learning. These books contain exercises and tutorials to improve your practical skills, at all levels!

Not satisfied with this document? We have related documents to Mitigating Catastrophic Forgetting in Reinforcement Learning, try searching with similar keywords: Mitigating Catastrophic Forgetting in Reinforcement Learning, Solutions Reinforcement And Answers Reinforcement, POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions, CANNABIS FORGETTING AND THE BOTANY OF DESIRE MICHA, Cannabis Forgetting And The Botany Of Desire Repost , Forgetting My First Real Kiss, Forgetting Sarah Marshall, Forgetting Your Sins What Does The Bible Say .pdf

You can download PDF versions of the user's guide, manuals and ebooks about Mitigating Catastrophic Forgetting in Reinforcement Learning, you can also find and download for free A free online manual (notices) with beginner and intermediate, Downloads Documentation, You can download PDF files (or DOC and PPT) about Mitigating Catastrophic Forgetting in Reinforcement Learning for free, but please respect copyrighted ebooks.