Blog

Notes on what we are working on and what we have found. Claims carry a number or a citation.

Published
28 Sep 2026
Author
Dan Lenton, unify
Read
30 min

Why we built Continual-ARC.

We ran three open agent harnesses through a stream of ARC tasks that return with fresh inputs, with their memory on and with it wiped before every instance. Memory saved them almost nothing. A floor that does nothing but keep each task's worked examples matched them, and a one-program-per-task library did far better.

Read the post

All posts

DatePostRead
Sep 2026
Why we built Continual-ARC.

We ran three open agent harnesses through a stream of ARC tasks that return with fresh inputs, with their memory on and with it wiped before every instance. Memory saved them almost nothing. A floor that does nothing but keep each task's worked examples matched them, and a one-program-per-task library did far better.

30 min
Sep 2026
The harness is not enough.

Harnesses are how today’s agents remember anything at all. They are a point-in-time fix, and continual parametric learning will replace them.

4 min
Aug 2026
Where the next run lives.

A scheduler has to write down when a recurring task runs next, and the obvious place is the task itself. We argue that each occurrence should instead be a row of its own, keyed by the facts that identify it, so that independent components converge on the same occurrence without a lock or a lease, and the definition holds only what a person asked for.

7 min
Jul 2026
Distilling repeated work into code.

An agent that meets the same task every hour pays each time to reason its way to the same steps, and may reason its way to different ones. We argue that once it has done such a task, it should distil it into a function whose control flow is fixed, with focused model calls only where the task needs judgement, and repair that function in place, after looking at the world, when the world changes.

8 min
Mar 2026
Demonstration without a mode.

Teaching an agent by showing it is usually built as a mode, with a capture pipeline and a store of learned skills that belong to it alone. We argue that it should fall out of three general properties instead, namely images that are ordinary message content, task sessions that stay open for corrections, and a library that any running task can write to.

6 min
Feb 2026
Tasks that outlive their first result.

A follow-up to a finished task usually starts a new task, which has lost the transcript and the working state that the first one built. We argue that a task which may receive further instructions should surface its result and then wait, with its transcript and its sandbox intact, so that a follow-up reaches the same loop in the same way as a correction.

7 min
Feb 2026
Functions and guidance, linked both ways.

Not everything an agent learns can be run. We argue that the procedures it learns should be kept as functions and the rules and walkthroughs as written guidance, in two libraries that are searched as peers and linked many-to-many, so that one rule can govern many functions and a function can be found without a document wrapped around it.

7 min
Feb 2026
Stateless calls, stateful agents.

A language model call remembers nothing of the call before it, and nothing can change it once it has started. We argue that an agent should supply what the call lacks with many loops on different clocks, a steerable handle on every piece of running work at every depth, and two kinds of memory kept outside the model.

8 min
Feb 2026
Forking the conversation, or starting cold.

A request that one model writes for another leaves out whatever seemed obvious inside its own conversation. We argue that wherever work is delegated, the delegating model should choose on each call whether the task forks the conversation, taking its history up to that point, or starts cold with the request alone.

9 min
Feb 2026
A fast brain and a slow brain for spoken agents.

A spoken reply has to begin within about a second, while a useful one can take a capable model ten seconds or more. We argue that a spoken agent should split the work between two models, a fast one that holds the conversation and never guesses, and a slow one with the whole context and every tool, which passes the fast one data but never words.

6 min
Jan 2026
Execution as a coordinate.

An agent that acts by writing code has to run it somewhere, and the two usual answers both fail over a long task, since a fresh sandbox forgets everything and a shared interpreter keeps too much. We argue that every execution should be a point the agent chooses, with what survives it, where it runs and which session it joins set on each call as independent arguments of one tool.

11 min
Jul 2025
Offline, hierarchical memory consolidation.

An agent that remembers has to keep deciding what is worth keeping from what it has just seen and done. We argue that this consolidation should run offline, away from the conversation, and that its summaries should form a hierarchy over both time and the number of interactions, with each level built only from the level below.

8 min
Jun 2025
Talking and acting in separate loops.

An agent that does real work for a person has to keep talking while it works. We argue that the conversation and the work should run as separate loops, with the work handed back to the conversation as something it can steer and question while it runs, and with the work able to ask the person for what it needs.

6 min
May 2025
Storing skills as functions.

An agent that works over a long period meets the same kinds of task again and again. We argue that what it learns about doing them should be kept as a library of Python functions rather than a list of notes. Functions call one another, so shared steps and shared abstractions each exist once, and every task that needs them calls them.

12 min