An agent should keep listening while it works
An agent that does real work for a person is often busy for minutes at a time, searching a website, filling in a form or drafting a message. The person does not stop talking while it works. They ask how it is going, add a detail they forgot, change their mind, or tell it to stop. The agent, for its part, often needs to ask them something before it can finish. In speech the problem is sharper, since a silence of a few seconds on a call already sounds like a dropped line.
Most agents run a single loop, in which the model reasons, calls a tool, reads the result and reasons again[1]. The person's next message is read only when the loop comes round to it.
A single loop has three limitations
Firstly, the person is blocked. A message sent while a tool is running waits until the current model call and every pending tool have returned, and a long tool can hold the conversation for minutes. Secondly, the work is opaque. While it runs, nothing can answer a question about its progress without stopping it. Finally, the work cannot ask. A tool that finds a detail missing can only fail or guess, since nothing connects a running tool back to the person. Gim et al. identify the first of these for single function calls, each of which blocks the model until it returns[3].
Talk and act in separate loops
We propose to run the conversation and the work as two loops. Our design gives the conversation to one part, which answers the person and never waits on the work, and gives planning and acting to another, with queues between them. A third part watches the conversation for requests and passes each one to the acting loop, either as a new task or as a change to the task already running. In the simplest version, each completed turn of speech is handed to a separate worker, and whatever the worker reports is spoken at once, cutting off anything the agent was in the middle of saying.
The two loops. The person talks only to the conversation loop, which passes requests, interjections and answers down to the acting loop, in blue, and receives progress, results and questions back, in coral.
Work comes back as a handle
Starting a piece of work returns a handle instead of a result. The handle can be given a new instruction, paused and resumed, stopped, and awaited for its result. The conversation loop is itself a tool loop, and each piece of work it has started adds tools for steering that one piece to the loop's next turn, named after the function they steer. During a search, the conversation can add an instruction to it, pause it, stop it, keep waiting for it, or answer a question it has asked, and it can ask the running work how it is going.
The tools that one running search can add to the conversation loop, each named after the function and the call it steers. Which ones are offered follows the search's state. Pause and resume alternate, clarify appears only once the search has asked a question, and every one of them goes when the search is done.
The layer between the conversation and the work offers two surfaces. One answers questions and changes nothing, and the other may change tasks, contacts or stored facts. Both return a handle, and both rebuild their tools on every turn from whatever is running at that moment. A question about progress reaches the running work through the first surface, and an instruction reaches it through the second.
The person always goes first
In the acting loop, a message from the person is taken before anything else, even while long tools are in flight. If the model is already partway through deciding its next step, the new message cancels that call, and the loop starts again with the message in view. The conversation so far is passed down to the work it starts. A task begun partway through a conversation then knows what the person has already said.
The work can ask
A running tool can put a question to its caller, and the question appears in the caller's loop as an unfinished tool call waiting for an answer. If the caller cannot answer, it asks its own caller, and the question rises until it reaches the person. The answer travels back down the same way, and the tool carries on from where it stopped. In one test, a request to email a friend about arriving at a barbecue reaches a tool that wants to know whether the sender is bringing anything. The question rises two levels to the person, and the answer comes back down to finish the email.
A question rising from the work to the person, in coral, and the answer passed back down, in blue. Neither loop above the tool has to stop for this to happen.
Related work
ReAct interleaves reasoning and actions in a single loop[1]. The Talker-Reasoner architecture splits an agent into a fast Talker, which converses, and a slower Reasoner, which plans, calls tools and writes the agent's state for the Talker to read[2]. Of the work we know, it is the closest to our design. We add control over running work in both directions. The conversation can steer the work as it runs, and the work can ask the conversation for what it needs. AsyncLM runs function calls concurrently with the model and interrupts the model when one of them returns[3], and LLMCompiler plans function calls to run in parallel[4]. Both keep a single loop. Moshi listens and speaks at the same time within one speech model[5], which addresses the same problem for speech itself, one level below the work.
Open questions
The layer between the loops owns at most one running plan at a time, and several running at once, each with its own steering tools, is untested. Every running piece of work also adds tools to the conversation, and a conversation with many tasks under way may be offered more tools than a model can choose among well.
The conversation loop is itself a model, which adds a model call to every turn. In speech it may need to be a faster model than the one doing the work. Deciding whether a new message is an instruction for the running work or a new request is left to that model, and we have not measured how often it chooses wrongly. A question that rises to the person also holds the work until it is answered, and a person who never answers leaves it waiting.
However, we argue that none of these weighs against separating the two loops. Each is a question of scale or of measurement, and we leave them to future work.
1. Yao et al. ReAct: synergizing reasoning and acting in language models. ICLR 2023.
2. Christakopoulou et al. Agents thinking fast and slow: a talker-reasoner architecture. ARXIV 2410.08328, 2024.
3. Gim et al. Asynchronous LLM function calling. ARXIV 2412.07017, 2024.
4. Kim et al. An LLM compiler for parallel function calling. ICML 2024.
5. Défossez et al. Moshi: a speech-text foundation model for real-time dialogue. ARXIV 2410.00037, 2024.