11 Jun 2025

Requests that carry their conversation.

A request leaves out what the conversation settled

An agent that talks in one loop and acts in another hands work across a boundary many times in one conversation. The conversation starts a task, the task calls a lookup that runs a loop of its own, and that lookup may call another. Each call carries a request written by a model, and the model writes it from inside a conversation. Whatever that conversation has already settled tends to stay out of the request, because to its writer it is obvious.

Suppose a person tells the agent that they enjoyed a phone call about basketball last week, and then asks what day that call was. The agent passes the question to a lookup over past messages and calls, worded as "What day was the conversation?". That week the person also exchanged emails with the same friend about a holiday. A lookup that reads only the request has two conversations to choose from and no way to tell which one is meant. A lookup that has also read the conversation can answer straight away, since the person said which call they meant two messages earlier.

the conversationpersonI really enjoyed our call about basketball last week.agentMe too.personWhat day was that call?ask("What day was the conversation?")the conversation,passed by the loopwith the conversationreads the conversation, then the requestanswers with the dayof the basketball callthe request alonefinds two conversations with Juliahas to ask which conversationis meant

The same request, handed on two ways. The lookup in aqua receives the conversation, in blue, along with the request, and answers with the day of the basketball call. The lookup in coral receives the request alone, finds two conversations with Julia that week, and has to ask which one is meant.

Handing work on has three limitations

Firstly, a tool call carries only the arguments its caller writes. An agent called as a tool in the OpenAI Agents SDK, for example, receives a single input string, and the caller's conversation is not passed to it[1]. The callee then meets an ambiguity that its caller could not see, and it can only guess or ask. A confident guess is hard to catch until it has done its damage, and a question interrupts the person about something they have already said.

Secondly, the usual remedy is a fuller request, and the model asked to write it is the one that left the detail out in the first place. However long the brief becomes, it is written from inside the conversation, and it omits whatever its writer did not think to mention.

Finally, the systems that do pass the conversation on do it by handing over control. In the same SDK, an agent that receives a handoff sees the whole earlier conversation, and it takes the conversation over from the agent that handed off[2]. That suits moving a person to a different specialist. A lookup is different, since it should answer one question and give control straight back.

The loop passes its conversation with every call

We propose that the loop making a call passes its own conversation down with it, and that the loop on the other side reads the conversation before the request. A tool receives the conversation if its function declares a parameter for it, and a tool that does not declare one is called exactly as before. Every lookup declares it and hands it to its own loop. Starting a task declares it too, and the task passes it on to the plan that carries the task out.

We give the inner loop the conversation as one system message, headed as broader context and marked read-only, with an instruction to resolve the next request in light of it. The request itself still arrives as the caller's first message. The inner model can then tell the task it has been given from the background it has been shown. The old messages arrive as data inside that one system message, not as turns of the inner exchange. An instruction the person once gave the outer agent then reads as background to the lookup, and not as an instruction addressed to it. A lookup's own description of the parameter calls it optional, read-only chat history.

A version of the basketball example is one of our tests. The question "What day was the conversation?" arrives with an outer exchange about enjoying a conversation on basketball, and the lookup returns the day of the basketball call.

The model never writes the conversation

We have the loop fill the parameter in, and we leave it out of the description the model reads, from its schema and from its docstring alike. A test keeps it hidden. In our first version every parameter of a tool went into its description, the conversation's included, and the model could see the parameter and write a value of its own. A model filling it in would paraphrase or shorten the history, which brings back the loss the parameter is there to prevent. It would also spend output tokens copying out something the loop already holds.

We then took it out of the description. Whether the conversation travels with a call is now a setting of the loop, on by default, chosen once by whoever builds the loop and not by the model on each call.

what the model readsask(text)Answer a question about pastmessages and calls.what the model writesask("What day was the conversation?")what the lookup is called withrequest"What day was the conversation?"contextthe conversation so far, added by the loop

What the model sees and writes, and what the call carries. The description the model reads has no place for the conversation, and the model writes the request alone. The loop adds the conversation, in blue, before the lookup runs.

Each handover adds a layer

A task that has received the conversation and then calls a lookup passes on more than it was given. Before the call, the loop combines the history it inherited with its own messages, leaving out the system message that carried the inherited part. It then attaches its own exchange beneath the last message it inherited. Nothing is repeated, and the history becomes a tree. A lookup started by a task, which the conversation started, finds the person's conversation at the root, with the task's own exchange hanging beneath the message that started the task. Its own request follows.

The same holds when the conversation asks a running task how it is going. A separate read-only loop answers the question, and it starts from a snapshot of the task's own exchange, which it receives as its context. The task itself is left untouched.

the conversationpasses its historya taskpasses its historya lookupwhat the lookup reads firstbroader context, read-onlyuserI really enjoyed our call about basketball…userSend Julia a note to say thanks.userThank Julia for last week's call.assistantask("What day was the call with Julia?")then its request"What day was the call with Julia?"

What a lookup two layers down reads before its request. The person's conversation, in periwinkle, is the root. The task's own exchange, in aqua, hangs beneath the message that started the task, and the lookup's request, in blue, comes after both.

Asking is for what the conversation leaves open

The conversation does not remove the need to ask. A lookup that can reach its caller is given a tool for asking, and the tool's description tells the model to use it whenever a request feels incomplete. In a second test, the same question arrives with no conversation at all, and the records hold both the basketball call and the emails about the holiday. The lookup asks which conversation is meant. Once told it is the one about basketball, it returns the same day as before.

The two tests together show the division we want. The conversation settles what the person has already said, and a question covers what nobody has said yet. Every question still interrupts the person, so a question that the conversation could have answered is a cost we would rather not pay.

What passing the conversation costs

Firstly, every call pays for the whole history, whether it needs it or not. A lookup that only wants a friend's email address may read forty messages to find it. Models also use long contexts unevenly, and recall what sits in the middle of one less reliably than what sits at either end[3]. Secondly, the history arrives as one block, and the inner model has to work out for itself which parts bear on its request. The header asks it to read the request in light of the history, but nothing marks the messages that matter. Finally, the context is a snapshot taken when the call starts. A task that runs for some minutes does not hear what the person says while it works, unless the conversation passes it on as an interjection.

The OpenAI Agents SDK gives an agent that receives a handoff the whole conversation by default[2], and an agent called as a tool only the input its caller writes[1]. AutoGen builds applications from agents that converse, and in a group chat every agent reads one shared conversation[4]. Anthropic describes an orchestrator model that breaks a task down and hands each part to a worker model[5], without settling what a worker sees beyond its part. Liu et al. show that models use information in the middle of a long context less reliably than information at its ends[3]. In the systems we know of, work that receives the conversation also takes over control, and work that hands control back receives only its request. We keep control with the caller, as a tool call does, and still pass the conversation down, one layer deeper at each handover.

Open questions

Firstly, the conversation goes with every call, and we are not sure that is the right default for every tool. A search of the web gains little from forty messages of history. A choice made per call, by the model that writes the request, could save most of that cost. We took the parameter out of the model's view precisely so that it would not make that choice, and we do not yet know which way is better. Secondly, we pass the history whole. A summary, or a selection of the messages relevant to the request, would be shorter, but a model would write it, and it could drop the very detail that the whole history was kept for. Finally, we have not measured how often the conversation answers a question that a lookup would otherwise have had to ask.

However, we argue that each of these is about how much of the conversation to send, and when. None of them makes a case for sending the request on its own.

1. OpenAI. Agents as tools. OpenAI Agents SDK documentation, 2025.

2. OpenAI. Handoffs. OpenAI Agents SDK documentation, 2025.

3. Liu et al. Lost in the middle: how language models use long contexts. TACL 2024.

4. Wu et al. AutoGen: enabling next-gen LLM applications via multi-agent conversation. ARXIV 2308.08155, 2023.

5. Anthropic. Building effective agents. Anthropic engineering blog, 2024.