A request is a summary of a conversation
An agent that talks in one loop and acts in another has to hand work across a boundary, and often more than one. The conversation starts a task, the task calls a lookup that runs its own loop, and that lookup may call another. At each boundary the caller knows things that its request does not say. The request is a summary of a conversation, written by a model that is inside the conversation, and the details it leaves out are the ones that seemed too obvious to write down.
Suppose a thread has touched two documents this week, a board deck and a weekly metrics report, and the person's contacts include two people called Sarah, one in finance and one in design. Earlier in the thread the person said that Sarah Chen in finance had asked for the metrics. They now say "send it to Sarah once it's done", and the conversation starts a task with the request "email the report to Sarah when it's finalised". A task that has read the conversation does not notice a question, since the answers are in what it read. A task that has read only the request cannot tell which report or which Sarah is meant. It has to ask, or to guess.
The same request, delegated two ways. The forked task, in aqua, reads the conversation up to the request, in blue, and sends the right report to the right Sarah. The cold task, in coral, reads the request alone and cannot tell which report or which Sarah is meant.
Existing ways of delegating have three limitations
Firstly, a task that starts from the request alone meets ambiguity that its caller could not see. Asking is safe but costs a round trip that interrupts the person for something the agent already knew. Guessing is worse, since a wrong assumption carried out confidently is hard to catch until it has done its damage. Systems that delegate this way ask the caller to compensate with a detailed brief. A multi-agent research system, for example, found that each subagent needs a detailed brief, with an objective and clear boundaries, and that vaguer briefs led subagents to duplicate each other's work[1]. However detailed, a brief is written from inside the conversation, and it leaves out what its writer did not think to mention.
Secondly, passing the whole conversation with every delegation makes every lookup pay for it. A web search does not need forty messages of history, and a model given a long context uses it unevenly, recovering information from the middle of it worse than from its ends[2]. Finally, where systems do pass the conversation on, the policy is fixed ahead of time by whoever wrote the system. A handoff in the OpenAI Agents SDK gives the receiving agent the entire conversation by default, and a developer can attach a filter to trim it[3]. Neither lets the delegating model decide, for the request in front of it, whether the history matters.
Fork or go in cold, chosen on each call
We propose that every tool which delegates work takes one more argument, a choice between two behaviours. With the argument set, the task forks the conversation. It receives the history up to the moment it was started, and then continues on its own. Without it, the task starts cold with only the request. The tool that starts a task forks by default, and its description tells the model to go in cold when the request says everything, as in a simple lookup or a web search.
The choice belongs to the model because it depends on the request and not on the tool. The same tool starts both the task that sends the report to Sarah and the task that searches the web for a train time. Every loop lets its model decide by default, and the argument is added to the schema of every tool that can take the conversation, with a description of when to pass it and when not to. The two other settings, always and never, remain for loops whose builder knows the answer in advance.
What the fork contains
A forked task finds the history in a section of its own instructions, which says that the request came from within a parent conversation. The history is there to explain the broader goal, and the task is told to focus on its own part of it. The messages keep their order, but their roles are renamed, and a message from the person becomes one from the "outer user". The task can then tell the parent's turns from the turns of its own exchange with its caller. The renaming also keeps an old message from the person from reading as an instruction addressed to the task now, and marks the history as context the system supplied, not text that a user slipped in.
The fork also leaves out what the task cannot use. The conversation's view of its running work lists the tools that steer each piece of it. A task that inherits that list can try to call them, writing code such as await stop_search_the_web_for__1() against a tool that does not exist in its own scope. We remove that section from the history before it is passed on, and keep the rest.
A fork, not a mirror
As with a fork in version control, the task receives the history at the branch point and not a live copy of everything that follows. The conversation carries on while the task runs, and some of what happens in it may matter to the task. When the conversation next steers the task, with an interjection, it sends only the part of the conversation that has changed since the task started, and the task receives it as a continuation of the history it already has. Each loop records what it has already forwarded to each of its inner tools, and no message is sent to the same task twice.
The choice made at the start sticks. A task started cold receives no history when it is later steered either, and a loop forwards updates only to those of its inner tools that forked. A web search started cold stays cold, however long the conversation grows while it runs.
A fork over time. The forked task, below in aqua, starts with the conversation's first three messages, and when it is steered later it receives only the four messages since, in blue. The cold task, above in coral, receives the request alone, and nothing when it is steered.
The same choice at every layer
A task delegates in turn, and its own model makes the same choice for each of its calls. When it forks, the history it passes on is the history it inherited with its own messages attached below the last of them. A lookup three layers down then receives the whole chain as a tree, with the person's conversation at the root and each layer's exchange beneath the message that started it. The same decision governs the calls that steer running work. A question to a running task starts a fresh loop and receives the full history, and an interjection receives only the continuation.
Code that the task writes does not have to carry any of this. A task often delegates from inside the code it runs, through calls such as a question to the agent's contacts. The sandbox wraps the object that exposes those capabilities in a proxy that adds the inherited history to every call that accepts it. The code makes the call as usual, and the history travels with it.
What a lookup three layers down receives when each layer forks. The person's conversation, in periwinkle, sits at the root, the task's own messages, in aqua, hang beneath the message that started it, and the lookup's request, in blue, arrives with both.
Related work
Liu et al. show that models use information in the middle of a long context less reliably than information at its ends[2]. Anthropic's multi-agent research system starts each subagent from a brief written by a lead agent, and finds that the quality of the brief decides whether subagents duplicate or miss work[1]. Cognition argues that agents should share full context and full traces, since actions carry implicit decisions that conflict when agents cannot see each other's reasoning[4]. The OpenAI Agents SDK passes the whole conversation on a handoff by default, with filters that a developer can attach to trim it[3]. In every system we know of, how much of the conversation a delegated task receives is fixed when the system is built. We give that choice to the delegating model on each call and at every layer of delegation, and the choice holds for every later call that steers the task.
Open questions
Firstly, forking is the default at every boundary, and we are less sure that it should be at the inner ones. A task deep in a chain may pass a long history to a cheap lookup that needs none of it, and the model making the choice cannot see how much it is passing. Secondly, the fork passes the history whole. A summary, or a selection of the messages relevant to the request, would be smaller, but either would be written by a model and would reintroduce the loss that the fork exists to avoid. Finally, a forked task learns of changes in the conversation only when it is steered. A detail that matters to a running task, mentioned when the conversation is busy steering something else, does not reach it. We have also not measured how often the model forks when it should, or what going cold saves.
However, we argue that these questions concern the defaults and the bookkeeping around the choice, and none of them favours taking the choice away from the model. Settling the defaults is left to future work.
1. Anthropic. How we built our multi-agent research system. Anthropic engineering blog, 2025.
2. Liu et al. Lost in the middle: how language models use long contexts. TACL 2024.
3. OpenAI. Handoffs. OpenAI Agents SDK documentation, 2025.
4. Yan. Don't build multi-agents. Cognition blog, 2025.