A model call is stateless and feedforward
A language model call takes text or images and returns something shaped, such as a JSON object or a program. Two of its properties matter here. It is stateless, in that the next call knows nothing of the last unless the surrounding software supplies the history again. It is feedforward, in that input goes in and output comes out, and nothing can change the computation halfway through. An agent that seems to remember was given the memory again in its prompt. An agent that accepts a correction in the middle of a task has usually cancelled one call and started another.
A single call is loosely analogous to one feedforward sweep through a cluster of neurons. The analogy does not go far, but it points at what is missing. A brain has specialised regions that run on different clocks, feedback paths that reach computation already in progress, and state that outlives any one sweep. Until models learn continually from the work they do, an agent has to supply these properties outside the model. Most of what makes an agent useful over a long task comes from how it does so.
A model call, on the left, keeps no state and cannot be changed once it starts. The agent on the right builds loops around such calls, in blue, with a path for steering running work, also in blue, and a path for questions back up, in coral. Its memory sits outside every loop.
A single loop supplies them poorly
The common design supplies all three with one loop, which calls the model, runs the tools it asks for, and appends the results to a growing transcript, in the pattern of ReAct[1]. This has three limitations. Firstly, presence and work share one clock. A message that arrives while the loop is busy has to wait in a queue, abort the work or be spliced into the next tool result, and nothing attends to the conversation while the work runs. Secondly, feedback reaches one level down at most. Text appended to the next tool result can steer the loop that reads it, but there is no handle on a loop three levels deep. Finally, the transcript is the only state. Working state and long-term memory both live in it, and it grows until it is truncated or summarised.
Regions on different clocks
We split the agent into loops that each own one kind of work and run at their own pace. Talking and acting happen in separate loops, and the agent keeps listening while it works. In speech the conversation splits again, into a fast model that holds the floor and a slow one that thinks. Away from any conversation, memory is consolidated in batches, with summaries that span anything from a day to a year. Each of these loops runs on its own clock, from a spoken reply within a second to a summary written once a week.
The model at the heart of the conversation loop is called as a stateless function over state that the system keeps. On each step the system renders the conversation and its running work into one message, and the model makes a single decision from it. Nothing the model needs is held inside the call, and a step that is cut short loses nothing that the next step cannot read again.
The loops of one agent on a logarithmic scale of time. Each runs at its own pace, from a spoken reply within a second to a weekly summary, and the model call at the bottom of every one of them, in blue, lasts seconds and keeps nothing.
Feedback at every depth
Every capability returns the same kind of handle, whether it is a task, a lookup in the agent's contacts or a web search. The handle has methods to ask about the work and to interject in it, and methods to pause, resume and stop it. Handles nest, since a task holds handles on the work it started in turn. A correction typed in the conversation reaches the task, and the task can pass it on to the search running inside it. In the other direction, a question from a loop several levels down rises to the person, and the reply is routed back to the call that asked it.
We made this part of the structure because feedback is hard to add afterwards. If steering means appending the person's words to the next tool result, it works one level down and nowhere else, since there is nothing below that to address. Much of the cost of having many loops comes from this, since every depth needs a loop of its own to address.
One handle type at every depth. A correction, in blue, passes down from the conversation through a task to the search running inside it, and a question from the search, in coral, rises the other way to the person.
What travels across a boundary
Every boundary between loops also raises the question of how much context should travel with a request. The delegating model chooses on each call whether the task forks the conversation or starts cold. It makes that choice again at each layer below.
Two kinds of memory
Statelessness splits into two needs that are easy to conflate. Working memory has to survive between calls within a task, and no longer. The agent keeps it in execution sessions, and each execution is a point the agent chooses, in a session it can name and inspect, and read without changing. Long-term memory is what the agent keeps after the work is done. We keep procedures as a library of functions that call one another, and the record of past interactions is consolidated offline into summaries. After every task a review reads the trajectory and decides whether anything in it deserves to join the library, which is often nothing. The review runs as a second loop after the task, and the task itself never sees the tools that write to the library.
Keeping the two apart lets each follow its own rules. A session can be thrown away when the task ends without losing anything the agent meant to keep, and the library changes only through a review that can see the whole of the work.
Where the analogy breaks
Feedback in a brain is continuous, with activity flowing both ways at once in the same tissue. Ours is discrete and turn-based, as a queue checked between steps or a generation cancelled and restarted. The difference shows as delay. Nothing here learns the way a brain does either. The weights do not move, and what we call consolidation writes rows to a database, which is a much shallower thing than the word suggests.
What the structure costs
Firstly, many loops are harder to read and to operate than one. A single loop can be followed from its first step to its last, and a fault in ours may lie in the path between two loops that each behave correctly. Secondly, the structure works against prompt caching, which makes a prompt that extends an earlier one much cheaper to read again. A conversation loop that renders its state afresh on every step caches only its system prompt. Finally, each loop reads its own context, and a history forked down a chain of delegations is paid for at every level it reaches.
Related work
Brooks builds the control of a mobile robot from layers that run concurrently, each a complete behaviour, with higher layers acting on lower ones[2]. Minsky describes a mind as a society of simple agents, none of which is intelligent on its own[3]. CoALA organises language agents by their memory, their actions and their decision procedure, and separates working memory from procedural, semantic and episodic long-term memory[4]. MemGPT manages a model's context as an operating system manages memory, moving information in and out of it[5]. Generative agents keep a stream of memories, reflect on it and plan from it[6]. ReAct interleaves reasoning and acting in a single loop[1]. Of the work we know, CoALA is the closest framework to ours. We build its separation of memories from loops that run concurrently on different clocks, and give every piece of running work the same steerable handle at every depth.
Open questions
The sharper objection is that all of this belongs in the model eventually. A model that holds a conversation while it reasons, keeps its own working state and accepts input partway through a computation would make much of this structure unnecessary. Moshi already moves the timing of speech inside a single model, listening and speaking at once[7], and we would remove our own handling of turns and interruptions once a model does it natively. Whether an agent should also distil repeated work into code is less certain, since a stored program trades the adaptability of reasoning afresh for an artifact that has to be repaired when the world changes.
However, we argue that the boundaries between loops will outlast whatever runs inside them. Even a natively interactive model has to be connected to a memory, to permissions and to many pieces of long-running work that it can steer. We leave the question of where each boundary should sit, as models improve, to future work.
1. Yao et al. ReAct: synergizing reasoning and acting in language models. ICLR 2023.
2. Brooks. A robust layered control system for a mobile robot. IEEE Journal of Robotics and Automation, 1986.
3. Minsky. The society of mind. Simon and Schuster, 1986.
4. Sumers et al. Cognitive architectures for language agents. TMLR 2024.
5. Packer et al. MemGPT: towards LLMs as operating systems. ARXIV 2310.08560, 2023.
6. Park et al. Generative agents: interactive simulacra of human behavior. UIST 2023.
7. Défossez et al. Moshi: a speech-text foundation model for real-time dialogue. ARXIV 2410.00037, 2024.