21 Jul 2025

Offline, hierarchical memory consolidation.

An agent that remembers has to keep deciding what is worth keeping from what it has just seen and done. We argue that this consolidation should run offline, away from the conversation, and that its summaries should form a hierarchy over both time and the number of interactions, with each level built only from the level below.

Every exchange leaves something to consolidate

An agent that works with people over months accumulates raw experience quickly. Each exchange adds messages to a transcript, and each action adds a record of what the agent did. Most of this is not worth keeping in full. Some of it changes what the agent knows about a person, some of it is a fact worth storing, and some of it creates or closes a task. Deciding which is which, and writing each into the right store, is consolidation. Over a long period the transcript grows too long to read, and what the agent can recall is whatever consolidation kept.

Most memory systems for language model agents consolidate online, inside the exchange that produced the material. MemGPT has the agent edit its own memory through function calls in the middle of a conversation[1]. Mem0 extracts and updates memories as each new pair of messages arrives[2]. A-MEM links a new note to related notes, and revises them, at the moment it is written[3]. Generative Agents writes reflections inside the simulation loop, whenever the importance of recent events passes a threshold[4]. We think consolidation belongs elsewhere.

Consolidating during the conversation has three costs

Firstly, it adds latency to the exchange. Every write is at least one more model call on the path a person is waiting on, and an agent in a live conversation has little time to spare between turns. Secondly, a single exchange is too little context. A remark that sounds like a fact may be corrected two messages later, and a passing detail may only prove important when it recurs. A consolidator that reads fifty messages at once can weigh recency against importance, keeping a job change mentioned earlier over small talk from a moment ago. Finally, it divides the agent's attention. The model doing the task must also decide what to remember, with the instructions and tools for each competing for the same context, and a failure while writing memory becomes a failure in the conversation.

Consolidate offline, in batches

We propose that consolidation runs as a separate process, away from the conversation. The agent logs every message as it goes, and every fifty messages the batch is handed to a consolidator that runs without blocking the exchange. The consolidator updates the stores concurrently. It finds new contacts and changed details, and refreshes each contact's biography and running summary. It mines the batch for facts worth keeping, and may restructure the tables that hold them. It also updates the task list. A failure in one update is caught without stopping the others, and none of them can reach the conversation.

The agent reads the results without paying for them. Each consolidated view is kept as a single, ready-made snapshot, and the planner places the latest one in its instructions at the cost of a lookup. Writing is slow and happens rarely, and reading is fast and happens on every turn.

online, while the person waitsa personmessagesthe agentreads the latest snapshotloggedthe transcriptevery messagea batch of 50consolidationruns on each batch, apartwritesthe storescontacts, facts, tasks, summariesread on the next turnoffline, off the conversation path

The two paths. Above, the exchange with a person, which reads the latest snapshot of the stores and never waits on consolidation. Below, in blue, consolidation of each batch of fifty messages into the stores.

This is loosely analogous to how memory is thought to consolidate in people. Complementary learning systems theory holds that the hippocampus records experience quickly and replays it offline, much of it during sleep. The neocortex can then integrate it slowly without overwriting what it already knows[5]. Lin et al. make a related argument for language models, and find that compute spent on a context before a query arrives cuts the compute needed at query time by about five times[6]. The analogy breaks in one respect. An agent has no night, and its offline periods have to be made, here by counting messages.

Summaries should form a hierarchy

Consolidation also has to condense what the agent has done. An agent should be able to say what it has been working on this week without rereading every action. We propose to keep these summaries in a hierarchy. At the base, a summary is written from the raw record of actions. Every level above is written only from the summaries of the level below, and each takes a fixed number of them. A week is summarised from seven days, four weeks from four weeks, twelve weeks from three four-week summaries, and fifty-two weeks from four twelve-week summaries.

Each summary then reads a small, fixed number of inputs, however long the history grows. The raw record is read once, at the base, and never again. Recent periods keep their detail while older ones keep only their gist, much as a person recalls yesterday in detail and last year in outline.

Each level is also triggered by the one below instead of by a clock of its own. When a summary is written, it announces itself, and a higher summary starts once enough of these announcements have arrived. A higher summary never starts before the inputs it needs exist.

by timethe raw record of actionspast day×7past week×4past 4 weeks×3past 12 weeks×4past 52 weeksby interactionthe raw record of actionslast interaction×10last 10 interactions×4last 40 interactions×3last 120 interactions×4last 520 interactions

The two hierarchies of summaries over the raw record of actions. Each rung is written only from the rung below, and the number beside an arrow is how many lower summaries make one higher summary.

Time alone is not enough

Activity is uneven. A person may send forty messages on Monday and none for the rest of the month. A summary of the past week is then crowded on one day and empty on another, and it says nothing about the last conversation once that conversation is more than a week old. We keep a second hierarchy alongside it, counted in interactions instead of days. Its rungs are the last interaction and the last 10, 40, 120 and 520, each again built only from the rung below. The time hierarchy says what happened recently, and the count hierarchy says what the most recent interactions were, however long ago they took place. Both are kept side by side, and an agent can be given either.

8 weeks agonowpast week: 2 interactionslast 10 interactions: 16 days

Interactions over eight weeks, one dot each, with two bursts. The past week, in aqua, holds only two of them. The last ten interactions, in coral, reach back sixteen days into the second burst.

MemoryBank condenses each day's conversations into a daily event summary, synthesises the daily summaries into a global summary, and lets memories fade along a forgetting curve[7]. Of the work we know, it is the closest to our time hierarchy, with two levels of summary where ours has five. Wang et al. summarise a dialogue recursively, writing each new memory from the previous memory and the next stretch of conversation[8]. Generative Agents builds trees of reflections, each more abstract than the observations and reflections beneath it[4]. RAPTOR builds a tree of summaries over documents, by clustering chunks by meaning and summarising each cluster[9]. MemoryOS moves dialogue from short-term to mid-term storage first in, first out, and from mid-term to long-term storage in segmented pages[10]. Among the works we have read, none consolidates along two axes, time and the number of interactions, or triggers each level from the completion of the level below.

Open questions

The window sizes are a guess. A day, a week and four, twelve and fifty-two weeks line up roughly with the calendar, and the interaction counts are ten times the numbers of weeks, but we have not tested other sizes. A fixed batch of fifty messages can also cut an exchange in half, and the stores lag the conversation by up to fifty messages. The transcript still holds everything said since the last batch, but the stores do not.

Errors compound upward. A summary is only as good as the ones beneath it, a mistake in one day's summary reaches every level above it, and the raw record is never reread to correct it. We also have no measure yet of what an agent recalls from these summaries, or of how much the count hierarchy adds to the time hierarchy.

However, we argue that none of these weighs against consolidating offline and in a hierarchy. Each is a question of tuning or measurement, and we leave them to future work.

1. Packer et al. MemGPT: towards LLMs as operating systems. ARXIV 2310.08560, 2023.

2. Chhikara et al. Mem0: building production-ready AI agents with scalable long-term memory. ARXIV 2504.19413, 2025.

3. Xu et al. A-MEM: agentic memory for LLM agents. ARXIV 2502.12110, 2025.

4. Park et al. Generative agents: interactive simulacra of human behavior. UIST 2023.

5. McClelland, McNaughton and O'Reilly. Why there are complementary learning systems in the hippocampus and neocortex. PSYCHOLOGICAL REVIEW, 1995.

6. Lin et al. Sleep-time compute: beyond inference scaling at test-time. ARXIV 2504.13171, 2025.

7. Zhong et al. MemoryBank: enhancing large language models with long-term memory. AAAI 2024.

8. Wang et al. Recursively summarizing enables long-term dialogue memory in large language models. ARXIV 2308.15022, 2023.

9. Sarthi et al. RAPTOR: recursive abstractive processing for tree-organized retrieval. ICLR 2024.

10. Kang et al. Memory OS of AI agent. ARXIV 2506.06326, 2025.