20 Feb 2026

Tasks that outlive their first result.

A follow-up to a finished task usually starts a new task, which has lost the transcript and the working state that the first one built. We argue that a task which may receive further instructions should surface its result and then wait, with its transcript and its sandbox intact, so that a follow-up reaches the same loop in the same way as a correction.

A follow-up usually starts from nothing

Suppose an agent has spent ten minutes producing a report on February's invoices. Along the way it found the right account, worked out that the billing API reports dates in local time, and built an authenticated client that it keeps in a variable. The person reads the report and says "now do the same for March". In most designs the task that produced the report has already ended, and the follow-up starts a new one. The new task has none of what the first one learned. It rediscovers the account and the quirk in the dates, slowly, or it gets one of them subtly wrong.

The alternative that keeps everything is to run the whole exchange as one supervised chat. The context survives, but the person has to follow every step, and the agent cannot go away and work for twenty minutes while the person does something else. Running the work in a loop of its own, apart from the conversation, solves the second problem, and a task that runs in its own loop can be corrected while it runs. What it does not settle is what happens to the task when it finishes.

Existing approaches have three limitations

Firstly, a follow-up that starts a new task from the request alone repeats the work of finding what the first task found. Secondly, a new task started with a summary of the old one depends entirely on the summary. A summary leaves out details that did not look important when it was written, and it is text, which cannot carry live objects such as an authenticated client or a loaded dataframe. Finally, keeping the whole exchange in one supervised loop keeps the context but ties the person to the work, since nothing runs while they are away.

Finishing is not ending

We propose that a task can be started as a persistent session. When it finishes a piece of work, it surfaces its response and then waits for the next instruction or for a request to stop. Its transcript stays in memory, and its sandbox stays open with whatever the work built up, since the sandbox is closed only when the session ends. "Now do March" wakes the same loop, which reads the new instruction after everything it already knows.

We did not add a separate operation for resuming a session. A follow-up reaches the session as an interjection, which joins the loop as a new turn from its caller, exactly as a correction does while the task is running. The caller never needs to know whether the task is mid-step or waiting, because the same call works in both states.

the person“Now do March.”one-shotFebruary reportendsfinds it all againMarch reportpersistentFebruary reportwaits, sandbox openan interjectionMarch reportwaits againstopped

The same follow-up against a task that ended and one that did not. The one-shot task, above, returns and is gone, and the follow-up starts a new task that has to find the account again, in coral. The persistent session, below in aqua, responds and waits, and the follow-up, in blue, wakes the same loop with its transcript and sandbox intact.

In the conversation, a persistent task stays among the running work after it responds. Its response is recorded as awaiting input, which is distinct from a progress update, since the task will not continue until it is told to. The conversation model sees the task with a note that it will not complete on its own, and that a response awaiting input means the task has finished its turn and needs the next instruction.

Deciding when to keep a task open

Keeping sessions alive becomes a decision, and we give it to the model that starts the task. Its instructions reduce it to one question, which is whether the person could plausibly send another instruction for this piece of work. A walkthrough, multi-step work that the person is watching, or a request framed as one step of a larger process should stay open. A standalone request that can be done in one pass, such as finding someone's email address, should not. Once a session is open, every later instruction that belongs to it goes to the session, and the model does not start a new task for each step.

The same instructions ask the model to combine objectives that belong together. A request to remember a procedure and a request to try it straight away share the same context, and splitting them into two tasks would lose the link between them. A persistent task never completes on its own, and the model has to stop it when the work is over.

Could the person plausibly send anotherinstruction for this piece of work?yesnoa persistent sessiona one-shot task“Walk me through the expense form.”“Draft the board update, I'll have comments.”“Start on the first step of the migration.”“What's Alice's email address?”“What's the weather in Berlin?”“Book the 9:40 train to Leeds.”

Requests sorted by the one question the model asks. Those after which the person could plausibly send another instruction, in aqua, start a persistent session. Standalone requests, in grey, run once and end.

What a session keeps that a summary cannot

A summary and a live session differ in more than their length. The session keeps every message of the task, including the ones that did not look important at the time, and it keeps the objects that the work created. A client that took three calls to authenticate, a dataframe that took four minutes to load and a variable holding the account id are all still in the sandbox. None of them can be written into a summary. The follow-up can use them directly, and it pays for none of the work that produced them.

a new task with a summarythe same sessionthe conversation so farone paragraphevery messagewhat the task learnedwhat its writer thought to mentionall of it, the date quirk includedlive objectsnonethe client, dataframe and account id

What a follow-up receives from a new task started with a summary, on the left, and from the same session, on the right. The summary keeps what its writer thought to mention, and loses every live object, in coral. The session keeps every message and every object, in aqua.

Hewitt's actors are persistent entities that process one message at a time from a mailbox and keep their state between messages[1]. A persistent session behaves as such an actor, whose messages are instructions. Jupyter keeps a kernel's state between the cells of a notebook[2], and the sandbox of a persistent session is kept in the same way between instructions. OpenHands lets a person keep talking to an agent inside the same sandbox, with the person present for each exchange[3]. MemGPT keeps an agent's memory across conversations by moving text in and out of its context, and it keeps text, not working objects[4]. Of the work we know, none lets a task that runs unsupervised in its own loop return a result and then wait for more, with its working state kept. We also know of none in which a follow-up arrives by the same call as a correction.

Open questions

Firstly, an open session holds resources. Each keeps its sandbox, and each holds one of the twenty slots that an agent has for running tasks until it is stopped. The wait for the next instruction has no time limit, and nothing closes a session that the person has abandoned. Secondly, the decision to keep a session open rests on the model's judgement of whether a follow-up is plausible, and it can be wrong in either direction. A one-shot task that does receive a follow-up loses its context, and a session that receives none holds its resources for nothing. Finally, a session that lives for a long time grows a long transcript. Every turn then costs more, and at some point the early context has to be summarised, which brings back the loss that the session exists to avoid.

However, each of these concerns how long a session should live, and we argue that none is a reason to end every task the moment it returns. We leave the right lifetime of a session to future work.

1. Hewitt et al. A universal modular ACTOR formalism for artificial intelligence. IJCAI 1973.

2. Kluyver et al. Jupyter Notebooks: a publishing format for reproducible computational workflows. ELPUB 2016.

3. Wang et al. OpenHands: an open platform for AI software developers as generalist agents. ICLR 2025.

4. Packer et al. MemGPT: towards LLMs as operating systems. ARXIV 2310.08560, 2023.