. 2025 Jan 8:18:1538741.

doi: 10.3389/fncom.2024.1538741. eCollection 2024.

Memory consolidation from a reinforcement learning perspective

Jong Won Lee¹, Min Whan Jung^{1

2}

Affiliations

¹ Center for Synaptic Brain Dysfunctions, Institute for Basic Science, Daejeon, Republic of Korea.
² Department of Biological Sciences, Korea Advanced Institute of Science and Technology, Daejeon, Republic of Korea.

PMID: 39845091
PMCID: PMC11751224
DOI: 10.3389/fncom.2024.1538741

Memory consolidation from a reinforcement learning perspective

Jong Won Lee et al. Front Comput Neurosci. 2025.

. 2025 Jan 8:18:1538741.

doi: 10.3389/fncom.2024.1538741. eCollection 2024.

Authors

Jong Won Lee¹, Min Whan Jung^{1

2}

Affiliations

¹ Center for Synaptic Brain Dysfunctions, Institute for Basic Science, Daejeon, Republic of Korea.
² Department of Biological Sciences, Korea Advanced Institute of Science and Technology, Daejeon, Republic of Korea.

PMID: 39845091
PMCID: PMC11751224
DOI: 10.3389/fncom.2024.1538741

Abstract

Memory consolidation refers to the process of converting temporary memories into long-lasting ones. It is widely accepted that new experiences are initially stored in the hippocampus as rapid associative memories, which then undergo a consolidation process to establish more permanent traces in other regions of the brain. Over the past two decades, studies in humans and animals have demonstrated that the hippocampus is crucial not only for memory but also for imagination and future planning, with the CA3 region playing a pivotal role in generating novel activity patterns. Additionally, a growing body of evidence indicates the involvement of the hippocampus, especially the CA1 region, in valuation processes. Based on these findings, we propose that the CA3 region of the hippocampus generates diverse activity patterns, while the CA1 region evaluates and reinforces those patterns most likely to maximize rewards. This framework closely parallels Dyna, a reinforcement learning algorithm introduced by Sutton in 1991. In Dyna, an agent performs offline simulations to supplement trial-and-error value learning, greatly accelerating the learning process. We suggest that memory consolidation might be viewed as a process of deriving optimal strategies based on simulations derived from limited experiences, rather than merely strengthening incidental memories. From this perspective, memory consolidation functions as a form of offline reinforcement learning, aimed at enhancing adaptive decision-making.

Keywords: CA1; CA3; dyna; imagination; offline learning; simulation-selection model; value.

PubMed Disclaimer

Conflict of interest statement

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Figures

**Figure 1**
Overview of the simulation-selection model. **(A)** Navigation sequences to two locations—one where a reward was obtained (high-value sequence) and one where it was not (low-value sequence)—are represented using different colors. Solid arrows indicate experienced sequences, while dashed arrows represent unexperienced (novel) sequences. **(B)** CA3 generates both experienced and novel (unexperienced) navigation sequences, independent of their value. Among these, CA1 selectively reinforces high-value sequences, whether experienced or novel. **(C)** The schematic diagram illustrates the basic circuit organization of CA3 and CA1. The numbers denote the average number of synapses for each projection pathway in a single CA3 or CA1 pyramidal neuron (Amaral et al., 1990). The extensive but individually weak recurrent collaterals in CA3 enable the generation of both remembered (experienced) and novel (unexperienced) sequences. In contrast, CA1, which lacks recurrent collateral projections but conveys strong value signals, selectively reinforces high-value sequences. Figure adapted from Jung et al., 2018, licensed under CC-BY 4.0.

See this image and copyright information in PMC

References

1. Addis D. R., Wong A. T., Schacter D. L. (2007). Remembering the past and imagining the future: common and distinct neural substrates during event construction and elaboration. Neuropsychologia 45, 1363–1377. doi: 10.1016/j.neuropsychologia.2006.10.016, PMID: - DOI - PMC - PubMed
1. Amaral D. G., Ishizuka N., Claiborne B. (1990). Chapter neurons, numbers and the hippocampal network. Prog. Brain Res. 83, 1–11. doi: 10.1016/S0079-6123(08)61237-6, PMID: - DOI - PubMed
1. Ambrogioni L., Ólafsdóttir H. F. (2023). Rethinking the hippocampal cognitive map as a meta-learning computational module. Trends Cogn. Sci. 27, 702–712. doi: 10.1016/j.tics.2023.05.011, PMID: - DOI - PubMed
1. Ambrose R. E., Pfeiffer B. E., Foster D. J. (2016). Reverse replay of hippocampal place cells is uniquely modulated by changing reward. Neuron 91, 1124–1136. doi: 10.1016/j.neuron.2016.07.047, PMID: - DOI - PMC - PubMed
1. Barron H. C., Auksztulewicz R., Friston K. (2020). Prediction and memory: a predictive coding account. Prog. Neurobiol. 192:101821. doi: 10.1016/j.pneurobio.2020.101821, PMID: - DOI - PMC - PubMed

LinkOut - more resources

Full Text Sources
- Frontiers Media SA
- PubMed Central
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Memory consolidation from a reinforcement learning perspective

Affiliations

Memory consolidation from a reinforcement learning perspective

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

References

LinkOut - more resources

Full Text Sources

Miscellaneous

Abstract

Conflict of interest statement

Figures

Similar articles

References

Related information

LinkOut - more resources

Full Text Sources

Miscellaneous