Code needs somewhere to run
Agents increasingly act by writing code, since one program can call several tools and branch on their results in a single step[1]. Much less attention goes to where that code runs, and to what survives it. Most systems give the model one executor with one fixed policy for state, chosen by whoever built the executor and not by the agent for the work in front of it.
Both of the usual policies fail over a long task. A fresh sandbox for every call forgets. An agent that spends four minutes loading a large export into a dataframe, prints one summary and loses the sandbox has to load the export again for its next question. One interpreter kept for good has the opposite problem. A variable from an abandoned attempt shadows a later one, and a rate that one task set is still set when the next task forgets to set its own. Notebooks show the same problem at scale. Of about 860,000 attempts to re-run public notebooks, only 24% ran without errors and only 4% reproduced their stored results, and the study names hidden state and cells run out of order among the causes[2].
The two usual executors. A fresh sandbox, on the left, closes after every call, and each new question pays again for the same four-minute load. One interpreter kept for good, on the right, lets a second task silently use a rate left over from the first, in coral.
Existing executors have three limitations
Firstly, an executor that starts fresh on every call throws away state that was expensive to build, and the agent pays for it again in time and in tokens. Secondly, a persistent interpreter lets one task's leftovers change the next, and nothing tells the agent which names are stale. Finally, systems that offer more than one place to run code give each its own tool with its own fixed policy. OpenHands, for example, gives its agents a bash shell and an IPython kernel as two separate actions, both kept alive for the whole task[3]. Neither action offers a clean run on request, or a look at the state that is guaranteed to leave it unchanged.
Every execution is a coordinate
We propose that one tool runs all of the agent's code, and that its arguments place each execution at a point in a small space with independent axes. The language is Python or a shell. The state mode is stateless, stateful or read-only. The session is picked by an id or by a name. The environment is either the agent's own interpreter or a dedicated virtual environment with its own packages. A session is identified by its language, its environment and its id together, and a Python session and a bash session with the same id are two sessions that know nothing of each other.
The three state modes mean the same thing wherever the code runs. A stateless call runs in a fresh namespace, a fresh subprocess or a fresh shell, and whatever it built is gone when it returns. A stateful call runs in a session that is kept. In a dedicated environment the session is a subprocess held open for it. In a shell it is a shell kept alive, whose working directory, variables, functions and aliases persist from one command to the next. A read-only call reads a session without being able to change it.
The state modes against the places code can run. Every cell is reachable from the same tool by changing its arguments, and read-only, in blue, is the mode neither usual executor offers.
The model chooses the coordinates from the work. The tool's instructions give it one line of reasoning for each mode. Stateful suits work that takes several steps, such as loading data and then analysing it, stateless suits a one-off check, and read-only suits a look that must not change anything. A single task might use four coordinates, with nothing set up in advance.
execute_code(language="bash", code="ls -lh exports/")
execute_code(language="python", state_mode="stateful", session_name="audit",
code="ledger = load_ledger('exports/ledger.parquet')")
execute_code(language="python", state_mode="read_only", session_name="audit",
code="ledger.pivot(index='account', columns='month')")
execute_code(language="python", state_mode="stateful", venv_id=3, session_id=1,
code="forecast = run_forecast('exports/monthly.csv')")
The four calls above. The shell command closes when it returns. The ledger is loaded once into a session named audit, in aqua, which is kept. The pivot runs against a copy of audit, in blue, which is thrown away, and the forecast runs in its own environment, whose session is also kept.
Reading without writing
A read-only call takes an existing session, copies its state into a throwaway sandbox, runs there, and discards everything the code wrote. In the agent's own interpreter the copy is of the session's globals. In a shell it is a snapshot of the working directory, variables, functions and aliases, restored into a fresh shell that is closed afterwards. In a dedicated environment the state is read from the session's process and passed to a one-shot process.
This mode makes an expensive session safe to explore. Suppose a session has spent four minutes building a dataframe, and the model wants to try a reshape it is unsure of. In the stateful mode a wrong guess can overwrite the dataframe, and the only way back is to build it again. In the read-only mode the model tries the reshape and reads the result, and the session it came from is untouched. The model no longer has to weigh the cost of rebuilding a session before it tries something it is unsure of.
Sessions the agent can name and inspect
Sessions have integer ids, but the tools prefer names. A name such as "audit" is easier for a model to keep track of across twenty intervening steps than an id such as 3. A registry maps each name to its session, and an agent may keep up to twenty sessions alive at once. One tool lists the sessions that exist and another shows what is inside one. The model can then check what still exists before it chooses where to run, instead of guessing from its own history.
The agent's plan for a task already runs in a Python sandbox of its own, and Python session 0 is that sandbox. The default stateful session is the plan's own namespace, and any other session is a deliberately separate one. This gives two layers of isolation, one between tasks and one between the sessions inside a task.
An environment for each function
The last axis exists because of a packaging problem that no amount of reasoning can solve. A function stored months ago may pin an old version of a library, and a function written last week may need the new version, whose interface changed. Both belong in the same library of functions. We let a stored function declare its own environment, whose specification is a project file held as data. The environment is built on first use and kept on disk, and it is built again only when its specification changes. For stateful work its processes are held open in a pool, one for each session. If a process dies, the pool starts a new one and retries the call once, though whatever state the old process held is gone.
Code inside a dedicated environment still calls the agent's capabilities, such as a lookup in its contacts, exactly as code anywhere else does. Each such call travels to the agent's process as a line of JSON, runs there, and its result comes back the same way. The environment isolates its packages without cutting the code off from the rest of the agent. We do not try to negotiate between environments. Two functions that need incompatible versions get two environments, and they never meet.
Two stored functions that pin incompatible versions of one library, each in its own environment. A call from either to the agent's capabilities crosses the boundary as a line of JSON, in blue, and runs in the agent's process, and the two environments share no packages.
Stored functions also take the same three modes as the agent's own code, called as .stateful(), .stateless() or .read_only() on the function itself, and the same holds for functions that run in the agent's own interpreter. A function from the library can then be tried against a session before it is allowed to change it.
Wrong combinations explain themselves
Some coordinates make no sense. A stateless call cannot name a session, and a read-only call must name one that exists. The tool returns these as structured errors and does not raise, and each error carries a suggestion that names what to change.
{
"error": "Cannot use state_mode='read_only' without specifying a session.",
"error_type": "validation",
"suggestion": "Provide session_id or session_name (must refer to an "
"existing session), or use state_mode='stateless'.",
}
The model can correct its next call from the suggestion alone, without having to read a traceback. The alternative to one tool is a tool for each combination, such as one for stateless Python and another for a persistent shell. That surface grows with the product of the options, the descriptions of the tools drift apart, and the model has to infer the differences between tools from their names. SWE-agent showed that the design of the interface between a model and a computer changes how well the model performs[4]. One tool with independent arguments asks the model about properties of the work, such as whether its results must persist and which packages it needs. We expect a model to reason about those properties more reliably than it picks from a growing list of executors.
Related work
CodeAct argues for executable Python as the actions of an agent, run in an interactive Python environment that keeps its state between turns[1]. TaskWeaver runs the code of each session in its own stateful Jupyter process, which keeps sessions from interfering with each other[5]. OpenHands gives its agents a persistent bash shell and a persistent IPython kernel inside one sandbox, as separate actions[3]. SWE-agent argues that the interface between a model and a computer should be designed for the model[4]. Nix installs every package with its own dependencies under a path derived from its inputs, which lets incompatible versions live side by side[6]. Of the agent systems we know, each fixes how state is kept for a given executor, and none offers a mode that reads a session without being able to change it. Our environments for each function apply the principle behind Nix to a library of functions that an agent writes for itself.
Open questions
The read-only mode is weaker than its name in three places. Firstly, the copy of an interpreter's globals is shallow. It stops any name from being rebound, but an object changed in place, such as a dataframe sorted in place, is changed in the session too. A deep copy would close the gap, at a cost in time and memory that grows with the state, and some objects, such as open connections, cannot be copied at all. Secondly, state that passes into a one-shot process in a dedicated environment has to be serialised, and only plain values such as numbers, strings and containers of them survive the trip. A dataframe in such a session is invisible to a read-only call. Finally, every copy covers the state of a process and nothing on disk, and a read-only call that deletes a file deletes it for good.
Beyond these, every session lives on the machine that runs the agent. A session on a machine across a network can die without notice, and it is not clear how to offer the same guarantees there. We test that a read-only call cannot rebind a name in its session, and the same for a shell's variables, but we have not measured how well a model chooses among the coordinates over a long task.
However, we argue that each of these is a limit of how one mode is implemented or measured, and none is a reason to take the choice away from the agent. We leave it to future work to close them.
1. Wang et al. Executable code actions elicit better LLM agents. ICML 2024.
2. Pimentel et al. A large-scale study about quality and reproducibility of Jupyter notebooks. MSR 2019.
3. Wang et al. OpenHands: an open platform for AI software developers as generalist agents. ICLR 2025.
4. Yang et al. SWE-agent: agent-computer interfaces enable automated software engineering. NEURIPS 2024.
5. Qiao et al. TaskWeaver: a code-first agent framework. ARXIV 2311.17541, 2023.
6. Dolstra et al. Nix: a safe and policy-free system for software deployment. LISA 2004.