31 Jul 2026

Distilling repeated work into code.

An agent that meets the same task every hour pays each time to reason its way to the same steps, and may reason its way to different ones. We argue that once it has done such a task, it should distil it into a function whose control flow is fixed, with focused model calls only where the task needs judgement, and repair that function in place, after looking at the world, when the world changes.

Repeated work is reasoned afresh every time

Much of the work people hand an agent recurs. An hourly triage of customer inquiries or a weekly report on last week's orders comes round again with fresh data and the same shape. An agent that reasons through such a task on every run pays for the whole derivation each time, often dozens of steps on a capable model. It also takes a fresh path each time, and a description such as "last week" can be read differently on two runs of the same task.

Storing what an agent learns as functions is the natural answer, but a recurring task is rarely pure code. Somewhere inside the triage is a judgement about what a customer wants, and somewhere inside the report is a sentence that summarises it. The question is how to keep the judgement and freeze everything else.

Reasoning afresh and freezing a script both fall short

Firstly, an agent that reasons afresh on every run pays for the full derivation every time, and may reach a different answer each time. Secondly, a script written once, with no model in it, cannot exercise judgement. It has to replace a judgement with rules, and rules for meaning are brittle. On the inquiries we use below, which are worded with vocabulary that crosses categories, a keyword classifier scores about 71% and a small model about 96%. Finally, a script with no model in the loop cannot notice when the world changes. When an API it reads renames a field, the script fails, or quietly produces the wrong answer, until a person notices.

Distil the task into code, with focused model calls

We propose that after an agent has done a recurring task once, the review that follows it asks whether the agent's reasoning was open-ended planning or bounded judgement inside a stable control flow. If the judgement was bounded, the review distils the trajectory into one function. Ordinary Python carries the control flow, the agent's own capabilities carry the side effects, and focused model calls with structured outputs, a low temperature and an explicit model carry the judgement. If the reasoning was open-ended, such as a search for new tools or the debugging of an unknown state, the task stays with a live agent.

The model calls go through a helper for one-shot queries from inside generated code. The agent's own instructions draw the line between the two kinds of substep. Exact lookups and arithmetic stay in code. Classifying and drafting, which turn on meaning, go to the model. A distilled triage has this shape.

async def triage_new_inquiries(base_url: str) -> int:
    last_seq = get_json(f"{base_url}/batches/last")["last_seq"]     # code
    inquiries = get_json(f"{base_url}/inquiries?after={last_seq}")  # code
    if not inquiries:
        return 0
    listing = "\n".join(f"{i['seq']}: {i['text']}" for i in inquiries)
    labels = await query_llm(                                       # judgement
        "Classify each inquiry as refund, bug, sales or other, "
        f"by what the customer needs.\n{listing}",
        response_format=InquiryLabels,
    )
    post_json(f"{base_url}/batches", route(inquiries, labels))      # code
    return len(inquiries)

A recurring task can then take the stored function as its entrypoint, and later runs call it directly, without an agent loop. The control flow does the same thing every time, and the model is asked only the question it is needed for, with a small prompt and a typed answer.

a live run, on every runreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasonactreasondozens of steps on a capable model,and a new path each timethe distilled functionread the cursorcodefetch new inquiriescodenone? stopcodeclassify them allone model callroute and file the batchcode

The same task, run live and distilled. A live run, on the left, alternates reasoning and acting for dozens of steps on every run. The distilled function, on the right, fixes the control flow in code and keeps a single focused model call, in blue, for the judgement.

Semantic downgrades are bugs

Distillation can fail in the opposite direction, by freezing too much. A review that sees twelve classified inquiries can write a function that classifies the next twelve with keywords, or drafts replies from templates copied from the examples it saw. The function passes on the examples and fails on the next week's inquiries. The review's instructions treat this as a bug. Where the trajectory interpreted or produced unstructured text, the stored function keeps that step as a model call with a stable contract, and it generalises by keeping the call, not by memorising the sample cases.

What distillation buys

We measured the triage task run hourly for eight runs of twelve inquiries each. The first run derived the workflow in six model calls and about 251,000 tokens, and the review after it stored the function. Each later run made one model call of about 645 tokens. All 96 inquiries were classified correctly, on the first run and on every run after it.

tokens per run, log scale1001k10k100k1M251,323run 1644run 2664run 3653run 4649run 5637run 6647run 7623run 8

Tokens per run of the hourly triage, on a logarithmic scale. The first run reasons its way through the task. Every later run, in blue, calls the stored function, which makes one model call of about 645 tokens, about 390 times fewer. Setting the task up and the review after the first run cost a further 1.2 million tokens once.

Repair in place, after looking

A stored function can break when the world changes, which is the reason a model has to stay within reach of it. A scheduled run of a stored function gets one repair attempt by default. When the function raises, a bounded repair loop runs and the same run retries. The repairing model sees the exception and the source. It also has a probe that runs a short read-only snippet in a fresh interpreter, to observe what the external interface returns now.

Its instructions separate what the function must preserve from what it may change. The outcome is the contract, meaning what it computes, which side effects it performs and where it delivers them. How it reads its inputs is not, since external interfaces change after a function is stored. The model is told to diagnose before it rewrites, to trust what it observes over the description the task was created with, and never to weaken the outcome to make an error go away. It rewrites the function in place, keeping its id, because the task that runs it refers to it by id.

We added the probe after the loop failed without it. Our first repair loop saw only the exception and the source, and the function's own validation messages describe what it assumed, not what the interface returned. When an order API renamed a field, that loop took four runs to converge and lost every one of them. With the probe, the repair looked at the live API, found the renamed field and fixed the function in one attempt, for about $0.18, within the run that failed. All ten runs delivered the correct batch, and the runs on either side of the repair made no model calls at all.

12345678910a field in the API is renamedwith the probewithout itone repair, $0.18lost while the repair works blind

Ten hourly runs, with a field renamed in the order API after the fourth. With the probe, above, the fifth run fails and is repaired in one attempt, in blue, and all ten runs deliver. Without it, below, the repair works blind and four runs are lost, in coral, before it converges.

DSPy expresses a task as a program of declarative modules that call a language model, and compiles the program by optimising its prompts[1]. AFlow searches over agentic workflows written as code, in which nodes call a model[2]. Self-Debugging has a model repair its own code from the result of running it[3]. Voyager refines a program with feedback from its environment until a check says the program works[4]. Robotic process automation records business workflows as scripts that replay a user's steps, and is known to break when the applications it drives change[5]. Of the work we know, DSPy's programs are closest in shape to ours. We derive such programs from an agent's own trajectory once a task has been done, and keep the model only where the task turns on meaning. When a program breaks, we repair it in place from what a read-only probe observes.

Open questions

Firstly, the probe is read-only by instruction and not by construction. It runs in an isolated interpreter with only the standard library, but nothing stops a snippet from writing to the API it was meant to read. Secondly, a model decides whether a task's reasoning was bounded enough to distil. A task distilled when it should have stayed with a live agent breaks in ways that a repair of its inputs cannot fix. Finally, our measurements are single runs of two tasks against seeded fixtures, with one pinned model. They show that the design can reach these numbers, not how often it does.

However, we argue that these are questions of how far to trust the probe and the review, and none of them argues for reasoning through a recurring task afresh on every run. How far each can be trusted is a measurement we leave to future work.

1. Khattab et al. DSPy: compiling declarative language model calls into self-improving pipelines. ICLR 2024.

2. Zhang et al. AFlow: automating agentic workflow generation. ICLR 2025.

3. Chen et al. Teaching large language models to self-debug. ICLR 2024.

4. Wang et al. Voyager: an open-ended embodied agent with large language models. ARXIV 2305.16291, 2023.

5. van der Aalst et al. Robotic process automation. Business and Information Systems Engineering, 2018.