Speech sets a deadline
People take turns in conversation with gaps of about a fifth of a second, across all ten languages in which this has been measured[1]. A reply that starts a second late already sounds hesitant, and a few seconds of silence sounds like a dropped line. Much of what a person asks an agent needs more than a reply, though. Finding an email address, checking a calendar or starting a task needs a model with the whole context and access to tools, and a capable model working this way takes ten to twenty seconds to respond.
Keeping the conversation and the work in separate loops leaves the conversation free while the work runs. In speech that is not enough, since a conversation loop run by the capable model is itself too slow.
One model cannot be both fast and informed
A single model faces three limitations here. Firstly, a capable model on every turn leaves the person waiting in silence while it thinks. Secondly, a fast model on its own lacks the context and the tools to answer, and when asked for a fact it tends to produce a plausible one. A phone number or a meeting time that sounds right is worse than none at all. Finally, a model that listens and speaks at the same time, such as Moshi[2], solves the timing of speech itself, but a reply that needs a lookup or a task still has to wait for them.
Two brains, one voice
We propose to split the agent's side of a spoken conversation between two models. A fast model speaks. It is a small model with its reasoning turned off, and its instructions carry little beyond who it is and who it is talking to, alongside the conversation so far. A slow model thinks, with the whole context and every tool, and passes the fast one what it finds. The person hears one voice. The split holds whether the fast model sits between speech recognition and speech synthesis, or is itself a model that takes in and produces speech.
The two models behind one voice. The fast brain, which the person hears, passes on what was said, and the slow brain answers with guidance, in blue, which carries data and never dialogue.
The slow brain passes data, never dialogue
The slow brain's guidance is limited to data. It may give the fast brain a fact, ask it to find something out from the person, or pass on a notification, such as a message that has just arrived. It may not steer the conversation or suggest what to say. We made this rule because two models directing one conversation pull it in two directions. The fast model has heard every word as it was said, and it is better placed to choose the words. It receives guidance as a notification that the person cannot see, and works it into what it says next.
The fast brain defers and never guesses
The fast brain follows two rules. It never states a fact, such as a phone number, a time or an amount, that has not already come up in the conversation. When asked for one, it says a single brief deferral, such as "Let me check on that", and ends its turn there. The answer arrives later as guidance, and the fast brain passes it on. It also never says that it cannot check, since from the person's side the agent can.
The first versions of these rules failed in both directions. The fast brain would defer and then guess anyway in the same turn, and it would defer on questions whose answers were already in the conversation. We now test deferral against both kinds of fast model, and test separately for deferrals that should not have happened.
A deferral paying off. The fast brain says it will check within a second and carries on with small talk, while the slow brain looks the answer up. The guidance, in blue, arrives about eleven seconds later, and the fast brain passes it on.
Guidance can arrive too late
In the ten or twenty seconds the slow brain spends thinking, the conversation moves on. The person may change the subject, or answer their own question. Guidance written for the old topic then arrives into a new one, and the agent says something out of place. We pass each piece of guidance through a fast check, which reads it against what has been said since the slow brain started thinking, and drops it if the conversation has moved past it.
For the same reason, rapid speech does not cancel the slow brain's work. People speak in fragments, and a slow brain restarted on every fragment would never finish. In speech the running step completes and only the waiting one is replaced, while in text each new message starts a fresh step with the latest context.
Guidance arriving too late. The person drops the question while the slow brain is still looking it up, and when the answer arrives the check drops it, in coral, instead of letting the agent read out a time nobody asked for any more.
Related work
Kahneman describes human thought as a fast, intuitive system working alongside a slow, deliberate one[3]. SwiftSage brings the split to agents for interactive tasks, pairing a fast module that acts with a slower one that plans[4]. The Talker-Reasoner architecture splits a conversational agent into a fast Talker and a slower Reasoner, which plans, uses tools and writes state for the Talker to read[5]. Qwen2.5-Omni splits a single multimodal model into a Thinker that writes text and a Talker that turns it into speech[6]. Moshi listens and speaks at once within one model[2]. Of the work we know, Talker-Reasoner is the closest to our design. We restrict what passes from the slow model to data, require the fast model to defer instead of guessing, and discard guidance once the conversation has moved past it.
Open questions
The relevance check is itself a model call. It adds a delay before guidance reaches the fast brain, and it can be wrong in both directions. Deferral also makes the agent say "let me check" often, and we do not know how often a person will tolerate it. The line between facts the fast brain may repeat and facts it must defer on is drawn in its instructions, and nothing enforces it. We have measured the design with tests of single turns, not with people in long conversations.
However, we argue that none of these weighs against splitting the two models. Each is a question of tuning or of measurement, and we leave them to future work.
1. Stivers et al. Universals and cultural variation in turn-taking in conversation. PNAS, 2009.
2. Défossez et al. Moshi: a speech-text foundation model for real-time dialogue. ARXIV 2410.00037, 2024.
3. Kahneman. Thinking, fast and slow. Farrar, Straus and Giroux, 2011.
4. Lin et al. SwiftSage: a generative agent with fast and slow thinking for complex interactive tasks. NEURIPS 2023.
5. Christakopoulou et al. Agents thinking fast and slow: a talker-reasoner architecture. ARXIV 2410.08328, 2024.
6. Xu et al. Qwen2.5-Omni technical report. ARXIV 2503.20215, 2025.