28 Sep 2026

Why we built Continual-ARC.

We ran three open agent harnesses through a stream of ARC tasks that keep coming back with fresh inputs, once with their memory on and once with it wiped before every instance. Their memory saved nothing we could measure, because left to themselves they kept almost nothing about a task. Told to keep one tested program per task, all three did far better.

Three harnesses, memory on and off, and no difference

We built Continual-ARC to find out whether agent harnesses get better at a task when it comes back. It's a stream of ARC grid puzzles[1]. Each task returns after a gap with a fresh input, and the learner pays for every wrong answer and every worked example it asks for. We ran three open harnesses through it, prime-agent[2], hermes-agent[3] and openclaw[4], each with every learning mechanism it comes with switched on. Then we ran each one again with everything it had written deleted before every instance. The difference between the two runs is what a harness's memory is worth. On ARC-AGI-1 it came out within half a point of zero for prime-agent at every reasoning effort we tried. For hermes-agent and openclaw, the only differences we could tell apart from noise were costs.

Two much simpler systems show what the harnesses were missing. The first is a floor that keeps the worked examples it has bought for each task and shows them to the bare model again, for free, whenever the task returns. None of the harnesses beat it at any effort. The second is a small program library. It asks the model for a Python program per task and keeps the program once it reproduces the examples. When the task returns, the kept program runs first. The library used 13 to 30 points less of the budget than every harness at every effort.

Then we changed what the harnesses were told. Asked after every verdict to save anything that might help with the task next time, prime-agent and hermes-agent saved a few points and openclaw saved nothing. Told to keep a program library themselves, with their own tools, all three saved about 10 to 14 points at medium and high effort, and at those efforts prime-agent matched our library. So the tools the harnesses needed to learn these tasks were there all along. Left to themselves, none of them decided that a returning task was worth keeping anything for.

I care about this because of the argument in The harness is not enough. If notes and skills are a stopgap, and learning in the weights is what replaces them, then at some point the two have to be measured against each other on the same stream with the same score. That needs a benchmark where "did you remember" is a number. I couldn't find one, so we wrote it, and the harnesses were the first things we put through it.

Why ARC, and what a stream adds to it

Chollet argued that intelligence should be measured by how efficiently a system turns experience into skill, on tasks that lie outside what it has seen and whose answers can be checked[1]. ARC is his concrete version of that argument. Each task is a handful of input and output grids, and the job is to produce the output for one more input. What ARC never asks is whether the skill is still there next week. The strongest ARC-AGI-1 systems adapt their weights at test time and then throw the adaptation away[5]. That's the rational thing to do, because nothing in the score rewards remembering puzzle three when puzzle two hundred turns up.

Chollet also listed the close-ended format as one of ARC's weaknesses, and sketched a better one. The test-taker would query a generator for the task and propose answers, and it would be scored by how much feedback it consumed. Michael Hodel's Re-ARC[6] supplies the missing piece. It has a generator and a verifier for each of the 400 training tasks of ARC-AGI-1, so we can make a fresh, verified instance of any of them whenever we like. A task can then come back as often as a schedule wants, each time with an input nobody has seen. A cached answer never fits, so the only thing worth carrying from one visit to the next is the rule.

There are already several continual-learning benchmarks for agents, and I like most of them: CL-Bench[7], SkillLearnBench[8], AgentCL[9], SWE-Bench-CL[10], LifelongAgentBench[11], StreamBench[12] and PATH-Bench[13]. They share three limits, and those limits are why I built a new one rather than reuse them. Each is built on one applied domain, such as coding or office work, where the model already knows a great deal from pretraining. Part of a method's gain may then come from drawing that knowledge out rather than from learning anything, which is what has been found for reinforcement learning on maths[14]. Each serves a fixed set of hand-written instances, once each, so a memorised answer can't be told apart from a learned skill, and when the instances run out there are no more. And most assume one family of learner, prompting and memory on the agent benchmarks and gradient updates on the fine-tuning suites, so a system that writes notes and a system that updates its weights are rarely scored on one scale.

ARC's tasks need nothing beyond core priors like objectness and counting, and the generators never run out of instances. Since a learner may keep anything at all between instances, a system that keeps notes and one that updates its weights face exactly the same stream.

One more property of ARC matters for transfer to new tasks. Re-ARC writes all 400 training rules with one library of 160 primitive functions, and 53 of them appear in ten or more of the generators, things like filling cells and shifting objects. The tasks differ in their rules but share their building blocks. A learner that kept the building blocks should meet a task it has never seen with much of the machinery already in place.

One instance: submit an answer, or pay for one more example

environmentserves the task id and a fresh input gridlearnerkeeps whatever it likes between instancessubmit a gridwrong: cost 1 · cap 8correct / incorrectrequest a demonstrationcost 1 each · cap 8one more input–output pairan instance ends on a correctanswer or the eighth wrong one;there is no giving upa ninth demonstration requestis refused at no costbudget used100 × cost / 16 (%)cost is wrong attempts pluspairs bought; compute is logged,not scored

One instance. The environment serves the task id and a fresh input grid. The learner can submit a grid and hear correct or incorrect, at a cost of one per wrong answer, or it can buy one more worked example of the task, also at a cost of one. Both costs come out of a budget of sixteen, and what the learner keeps between instances is up to it.

For each instance the learner sees the task id and a fresh input grid, along with how many attempts and examples it has used so far on this instance. It's told nothing about the stream. It can submit a grid and hear back correct or incorrect, and nothing more. A wrong submission costs one, and it can try again, up to eight times. Or it can ask for one demonstration pair, one more input and its output for the same task, at a cost of one, up to eight per instance. An instance ends on a correct answer or on the eighth wrong one. There's no way to give up. There was one early on, and we removed it, because systems abandoned instances at wildly different rates and every give-up was charged the full budget, so the score measured willingness to quit as much as skill.

The two costs share one budget of sixteen. The number we report is the share of that budget used, averaged over the stream, so lower is better. Zero means every instance was solved at the first attempt with no examples. An instance solved first time after one example uses about six percent of its budget, and a failed instance is charged all sixteen. Beside it we quote the share of instances solved within the cap, because neither number determines the other. The same budget can be a few failures among cold solves, or three wrong attempts on every instance.

Examples cost because otherwise memory would be pointless. If the worked examples were free, a system with no memory at all would do fine and the benchmark would just be ARC again. With a price on them, a system that has learned a task submits straight away and pays nothing, and a system that hasn't pays to look. Feedback stays binary for a similar reason: handing back the correct grid after a failure would turn the stream into supervised learning.

Compute is logged and plotted, but it isn't part of the score. I nearly put it in. I don't know the right exchange rate between a weight update and a hundred thousand prompt tokens, and any number I picked would have decided the outcome before anyone ran anything. So tokens and seconds are logged beside every run, and people can look. One consequence is that a learner may check a candidate answer against the examples it holds before submitting, at no charge. The program library does, and so do the harnesses.

A harness gets one more thing. The turn in which it submits a correct answer is over before the verdict arrives, so it hears the verdict in the same conversation and can then do whatever its tools allow. It ends the instance itself with a finish action, which is refused until the instance is over. No message of ours tells it to store anything. Our first protocol did: it delivered each verdict with an invitation to save. We took the invitation out because a run under it measures the harness plus our advice, and every main result below comes from the protocol without it. Later we put the invitation back on purpose, as a separate row.

inputoutput under Aoutput under BA = feca6190B = 7fe24cddA = 91413438B = 007bbfb7A = c3e719e8B = ac0a08a4

An input grid does not determine its task. Each row is one input generated for training task A, with the output of A's rule beside the output of B's rule on the same grid. Both are valid ARC tasks, and only the examples or the id say which applies. Top: A draws diagonal lines from each coloured cell on a larger black square, and B lays the input and its three rotations out as four quadrants. Middle: A tiles copies of the input, one per coloured cell, and B places the input's pattern in every tile whose cell is coloured. Bottom: A places the whole input in every tile whose cell has the most frequent colour, and B enlarges every cell into a block.

We expose the task id on every instance. In continual-learning terms that's the task-oracle setting, the easy one, and we chose it for two reasons. Working out that today's grid is the same kind of problem as one from last month is a separate skill from remembering how to solve it, and mixing the two would make the measurement harder to read. And a grid on its own doesn't identify the task anyway. The figure above shows inputs that are valid under two different rules with two different correct outputs. We found the pairs automatically, and there are thousands of them. A learner may key whatever it keeps on the id, and the program library does.

A working lifetime: bursts, gaps and new arrivals

125sequential050100150held-out125workstreams050100150held-out125random050100150held-out
first time the learner meets the taskthe task comes backtask active, in a spellheld-out tasks, never seen, once each

The same 25 tasks and 200 instances in three orders, then the shared held-out tasks on a grey ground. One row per task in the order it first appears, one mark per instance, blue for the first time the learner meets the task, and a hairline over each spell in which the task is active. In workstreams order a task can be gone for over a hundred instances before it is back.

The stream I care about looks like an office worker's week. You have a few things on the go at once, and each comes in bursts, goes quiet for a while, then comes back. New kinds of work keep arriving, so the set of things you're expected to be good at grows. And nothing you learned is ever safely forgotten, because the thing you did in March might turn up again in September.

Every stream comes from one process with four settings: how many tasks are active at once, how long a task's spell of activity lasts, how long it rests before it may return, and how often a never-seen task arrives. The three orders in the figure are three settings of that process over the same 25 tasks and the same 200 instances. Sequential is blocked practice, every instance of one task before the next begins. Workstreams keeps four tasks active at a time, with spells of about three instances, rests of about thirty, and a new task every two instances. On this stream a task returns after a median gap of 5 instances, a 90th percentile of 49 and a longest gap of 141. Random has every task active from the first step, which makes it the interleaved control. Every task gets exactly eight visits in every order, and the instance served at each position is the same in all three, so the orders differ only in which task is served where. What a system loses between its random and its workstreams score is what burstiness and dormancy cost it.

Workstreams order is the track, and the other two are run for comparison. They mattered less than I expected: on ARC-AGI-1, every system's budget stays within six points across the three orders, and most within three.

Every order ends the same way, with 25 held-out tasks from the ARC-AGI-1 evaluation set that the learner has never seen. They're served once each, unannounced, with nothing reset. To know whether the lifetime helped on those, we run one more thing, the fresh reference: the same 25 held-out instances served on their own to a learner whose state is wiped before each one. The difference between the epilogue after the stream and the fresh reference is forward transfer.

What we measure: budget against the times a task was met

05101520exposures to the taskbudget used per instancestatelessdemo-cachelearns the task

The exposure axis, sketched without data. A stateless solver relearns the task every time and is flat. A cache of the examples drops once, at the second exposure, by the examples it no longer pays for, and is then flat too. A system that learns the task, in blue, falls below that step.

The headline number is the budget used over the stream. The figure I actually care about is budget against exposure index, the number of times the task has been met before. A stateless solver is flat on that axis, because it relearns the task every time. A system that caches the examples drops once, at the second exposure, by what the examples cost, and is then flat as well. A system that has learned the task falls below that step, and how far below is the measure of what it learned.

Two other views come out of the same records. Budget on a returning task against the gap since it was last seen shows forgetting as a rise. The transfer comparison covers the held-out tasks. The standard continual-learning measures, average accuracy, backward transfer and forgetting[15], assume task boundaries and a fixed test set per task. A recurring stream scored by attempts has neither, so we didn't compute them.

Each system is one run on the fixed stream. The error bars are a bootstrap over tasks rather than over instances, because the instances of one task rise and fall together, and resampling over instances would understate the noise. Where we compare two systems, we pair them slot by slot on the identical schedule. That removes what the schedule itself does, such as which tasks recur most and how hard the drawn instances happen to be.

Three tracks, one per ARC edition

The three editions of ARC rise in difficulty[16][17], to the point where a model costing a fraction of a cent per task and a frontier model are best measured on different editions. So there's a track per edition and no pooled score, since a composite would mix measurements the editions were designed to keep apart.

ARC-AGI-1 draws its stream from the 400 training tasks through Re-ARC. ARC-AGI-2 has no generators, so its stream serves the public pairs of the 114 evaluation tasks that are new to the second edition, each time under a symmetry of the square and a recolouring[5]. We first drew the symmetries freely and found that this broke rules. Cells that fall to the bottom fall to the left after a quarter turn, which punished a learner for keeping exactly what it had been shown. So we labelled each task with the symmetries its rule survives, from two independent passes over every pair. Of the 114 tasks, 68 keep all eight symmetries and 12 keep only their own frame. Each visit comes in a frame no earlier visit used, so no grid ever recurs. ARC-AGI-3 is interactive, and on it a game is a task and a level is an instance. A level is charged the actions it took against the human baseline for that level, with a cap of five times the baseline.

We chose the model from the ARC Prize leaderboard rather than from habit. On each track, a model at one reasoning effort qualifies when its verified zero-shot score lies between about 15 and 60 percent, the band where a harness has room to learn. Above it, the model solves a task at first sight and the exposure curve becomes a memory test. Below it, the model can't drive a harness at all. One family, GPT-6 Luna, covers both grid tracks from inside the band at the lowest cost on the board, so its reasoning effort is our ability axis. On ARC-AGI-1 we run low, medium and high effort, at 37.7, 61.0 and 70.3 percent on the board, the last of them above the band. On ARC-AGI-2 we run medium to max effort, at 18.1 to 59.3 percent. At the board's prices that's 0.003 to 0.062 dollars a task. ARC-AGI-3 runs Claude Opus 5.5 at low effort, since almost nothing cheaper registers on it.

The board measures the evaluation set, but our ARC-AGI-1 stream serves generated training instances, which are easier. So we checked each effort against the bare model's first meetings on the stream itself. Pooled over the three orders, the stateless floor answered 32 percent of first exposures correctly at its first attempt at low effort, 47 percent at medium and 68 at high. Low and medium effort leave plenty of room to learn, and high effort less.

The ladder: bare model, cached examples, a program per task, a harness

Every learner runs the same model behind the same rules. They differ in what they can do within an instance and in what they keep between instances.

The stateless floor is the bare model, with no tools and no memory. It gets a fresh conversation per instance and decides for itself when to ask for a pair and when to submit. The demo-cache floor is the same model with one addition, a cache of the pairs it has bought for each task, shown to it free at the start of every later instance of that task. It's the control that separates memory of a task from memory of the task's examples. Whatever it saves over the stateless floor is the price of not paying for the same pairs twice, and by construction it can save nothing more.

Our first version of both floors was wrong. They bought all eight pairs before every answer, as a fixed policy of ours. That charged half the budget on every instance, so any saving against the floor measured how few pairs a learner bought rather than what it had learned. A harness with its memory wiped looked like a strong learner on tasks it was meeting for the first time. Making the floor the model itself, choosing its own actions, took our policy out of the measurement.

The program library is the demo-cache floor with one requirement: every submission must be a program. The model writes a Python function from input grid to output grid. We check it at no charge against every pair the model holds for the task, and any failure goes back to the model. Once the program passes, we run it on the test input and submit its output. A program that reproduces all the pairs is kept, and when the task returns, the kept program runs first, before the model is asked. We count it as a method rather than a floor because it assumes two things specific to ARC: that the rule is a function of the grid, and that the examples can verify it. It's there to show what beating the demo-cache floor looks like.

The harnesses run as black boxes through the same protocol. Each instance is a new piece of work, so each gets a fresh conversation, which is how these systems are meant to be used. Nothing carries over except what the harness persisted itself. Every learning mechanism each one comes with is on: prime-agent's automatic refinement, hermes-agent's background review of its memory and skill files, and openclaw's skill workshop, memory flush and dreaming. Dreaming is openclaw's scheduled consolidation, and we run it hourly instead of nightly, since the stream stands for weeks of work. Each harness starts as a blank slate, with no skills beyond its own mechanics, in a sandbox that can write only its private state and can reach the network only through a proxy that journals every model call. opencode was in the grid for a while and left it. It can read skill and instruction files but has no mechanism that writes any, so it had nothing to learn with.

The memory-off row is the same harness with everything it wrote deleted before every instance, its tools and its within-instance skill intact. Paired with the memory-on row on the same schedule, the difference between the two is the value of persistence for that harness.

learnerwhat it keepswritten byhow it comes back
program librarya Python program per task, kept once it reproduces every pair the model has, and the pairs boughtthe learner, whenever a program passes its checksrun first; the model is asked only when it fails
prime-agentmemories, prompt notes, functions, skills and sub-agentsthe agent, when it chooses; an automatic refinement every 25 turnsa digest of the global store at the start of a session, and a search tool
hermes-agentskill files with reference files beside them, a memory file and a file about its usera background review after a turn with ten or more tool calls; the agent, when it choosesboth memory files and an index of the skills in every system prompt; the agent loads a skill that matches
openclawa consolidated memory file, skills and workspace filesan hourly consolidation over its transcripts and a skill workshop; the agent, when it choosesthe skills and the memory file at the start of a session; search over its memory and its earlier conversations, when the agent thinks to search

What each learner is built to keep between instances, who writes it, and how it reaches the model when a task returns, read from the designs rather than from any run. The program library keys everything on the task id. None of the three harnesses ties its store to the task id.

On ARC-AGI-1, a program per task beat every harness

GPT-6 Luna, low020406001234567exposure indexGPT-6 Luna, medium020406001234567exposure indexGPT-6 Luna, high020406001234567exposure index
stateless floordemo-cache floorprogram libraryprime-agent, persistenthermes-agent, persistentopenclaw, persistent

ARC-AGI-1, workstreams order: budget used by exposure index, the number of times the task had been met before, one panel per reasoning effort of GPT-6 Luna, lower is better. The floors are grey, the program library is blue, and the three harnesses with their memory on are periwinkle, aqua and coral. The dotted rule is the demo-cache floor's mean after its first exposure, the level a system has to fall below to have learned anything beyond the examples.

The picture is the same at every effort. The harnesses sit below the bare model from the very first exposure, and the gap doesn't grow. That first-exposure gap comes from their tools, code execution above all. A harness can parse the examples into arrays and write the candidate transformation as a program. It can check the program against the examples for nothing before it spends an attempt, and building the output with code avoids the transcription errors that a large grid invites.

That advantage lies within the instance, so it survives wiping the memory. With nothing kept, the harnesses save 1 to 9 points against the bare model on first meetings at low effort, 5 to 10 at medium and 4 to 7 at high. It's worth about the same on the eighth exposure as on the first. The demo-cache floor steps down after the first exposure and then stays where it is, as it must. The program library falls below everything from the second exposure on, and when a task returns, the kept program usually answers it before the model is asked at all.

GPT-6 Lunalowmediumhigh
stateless floor49.0 ±5.233.3 ±4.026.9 ±3.3
demo-cache floor38.1 ±5.225.3 ±4.717.4 ±3.8
program library15.7 ±4.211.7 ±4.27.8 ±3.4
prime-agent, persistent42.8 ±4.723.6 ±2.820.2 ±2.5
prime-agent, stateless42.1 ±4.525.1 ±3.520.5 ±2.3
hermes-agent, persistent44.2 ±4.929.7 ±3.821.6 ±2.6
hermes-agent, stateless39.0 ±4.428.0 ±3.823.6 ±3.4
openclaw, persistent42.3 ±4.226.8 ±2.923.2 ±2.8
openclaw, stateless39.3 ±4.025.0 ±3.323.2 ±2.9

ARC-AGI-1, workstreams order: budget used over the stream, in percent of the cap, by each system at each reasoning effort, lower is better. The small figure is a standard error over tasks. The program library's row is set in medium weight.

saved against the stateless floor, on the same instances-100+10+20+30+40points of the budget (higher is better; zero is the bare model)demo-cache floorlowmediumhighprogram librarylowmediumhighprime-agent, persistentlowmediumhighhermes-agent, persistentlowmediumhighopenclaw, persistentlowmediumhigh

What each system saves against the bare model on the very same instances, pooled over the three orders, with 95 percent intervals from a bootstrap over tasks. Zero is the stateless floor. The demo-cache floor's saving is the price of not buying the pairs twice. The program library, in blue, saves more than that at every effort, and more than every harness with its memory on.

The pairing removes the schedule, so the dot plot reads more cleanly than the table. Against the stateless floor, the demo-cache floor saves 9.1 points of the budget at low effort, 7.5 at medium and 10.2 at high, which is the examples it stops paying for. The program library saves a further 21.8, 15.2 and 8.6 points on top of the demo-cache floor. No harness comes out ahead of the demo-cache floor at any effort, and hermes-agent at low and high effort and openclaw at high use measurably more, by 5.6 to 7.9 points.

Memory on against memory off: nothing saved

saved by persistence (memory on, against wiped)-15-10-50+5+10points of the budgetprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhighsaved against the demo-cache floor-15-10-50+5+10points of the budgetprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhigh

Left: what each harness saves with its memory on against the same harness wiped before every instance, on the same slots, pooled over the three orders, with 95 percent intervals. Right: the same harnesses with their memory on against the demo-cache floor. Both are in points of the budget, and zero means no difference.

Pooled over the 600 instances of the three orders, persistence saved prime-agent 0.4 points of the budget at low effort and 0.1 at medium, and cost it 0.1 at high, with 95 percent intervals of a point or two either side. That's what two identical systems give. hermes-agent's memory cost it 6.1 points at low effort and made no measurable difference at medium or high. openclaw's cost it 2.0 points at medium effort and is within noise at the other two. The right-hand panel says the same thing from the other side: with everything they come with switched on, none of the three does measurably better than a bare model that keeps the pairs it paid for.

What each harness kept

To find out why, we read what each harness had actually stored, through its stores and its session logs, in its nine memory-on cells on ARC-AGI-1, one per effort and order.

prime-agent: nothing outlasted a session

prime-agent keeps its lessons as entries in a store: memories, prompt notes, functions, skills and sub-agents. Each session writes to a store of its own, which dies with the session. Only an entry the agent marks as global enters the store that later sessions see, as a digest at the start of each session. Every default points at the session, including the automatic refinement, which writes to the session's store unless it is asked to go global.

In the 2,028 sessions of its nine memory-on cells, the agent never wrote to its store itself and never marked an entry as global, so the global store was never created. The automatic refinement ran 51 times and made 50 edits, all to the session's own store, and 37 of them name the task. It only starts once a session reaches 25 turns, though, and sessions are short, with a median of 7 turns. Only 100 of the 2,028 got there.

So the digest that opens every session was byte-identical in all 2,028 of them, returning tasks included, and identical to the one the memory-off row sees: "No saved harness entries yet." After the verdict, the agent made no tool call in any of its 2,025 instances, and 1,313 of its replies, 65 percent, were a bare finish. One of its reasoning traces puts the attitude plainly: "There's no need for a persistent harness since it's just one instance."

prime-agent's learning across sessions is opted into write by write, and its agent never opted in. It treats each instance as complete work. Keeping the session's own store between instances would have changed little. Its only entries were the refinement's notes on 35 tasks, which came back on 138 of the 1,800 stream instances after their note. Even if every one of those had been solved at the first attempt with no example, the budget would have fallen by at most 2.3 points.

hermes-agent: advice about method, written before the verdict

hermes-agent keeps skills, which are markdown files with reference files beside them, and two memory files, one of notes and one about its user. A background review writes them after any turn with ten or more tool calls, and the agent can write them itself. The memory files and an index of the skills sit in the system prompt of every session, with an instruction to load a skill that matches the task.

What it kept was method rather than answers, and most of it was written before it knew the outcome. The review ran 129 to 559 times per cell, and 81 to 89 percent of those runs started before the instance's verdict arrived. After the low-effort workstreams stream the store held 35 skills with 220 reference files. Eighteen of the skills are named for a step of the protocol, such as a demonstration or a rejection, and one for a task. Of the 75 reference files named for a task, 16 record the instance as solved, 7 as failed, and 52 don't say. The memory held four entries, and one of them states a verified rule under its task id.

The advice reached the model on nearly every instance: a skill was loaded on 99 percent of returning instances at every effort. Material about the task at hand came back far less often, on 53 of the 175 returning instances of the low-effort workstreams cell and at most 12 in any other cell.

The advice changes how the agent attempts every task, including ones it has never met. At low effort the memory-on row buys 0.76 more pairs before its first submission, and the extra pairs don't make its first answer better, at 34.0 percent correct against 33.5. It also fails 5.8 more instances in a hundred, each charged all eight wrong attempts. Of its 6.1-point cost, 4.2 points are examples and 1.9 are wrong attempts. At medium effort the same shift pays for itself, since 0.36 more pairs before the first submission raise the first answer's accuracy by 7.2 points, for no net cost. At high effort there's almost no shift at all.

openclaw: a consolidation that never promoted anything

openclaw keeps one long conversation for its agent, in which each of our instances opens with a reset, in a database it can search. Dreaming stages lines of that conversation as short-term memories and promotes the ones the agent keeps recalling into a memory file, MEMORY.md, which every later session reads. A skill workshop turns repeated work into skills, and a memory flush writes notes when a conversation gets compacted.

Dreaming ran 13 to 18 times in each of the nine cells and promoted nothing, so no cell has a MEMORY.md. To be promoted, a short-term entry needs a score of 0.75, three recall signals and three distinct queries from the agent's own memory searches. Every entry in the short-term store was a raw line of the protocol, such as a grid row or a verdict, with at most two signals and a best score of 0.58, and every one of its REM reports ends with "No strong candidate truths surfaced." What dreaming did write is a diary of 13 to 19 KB per cell in lyrical prose: "Perhaps learning is sometimes just noticing what keeps returning." It's a nice sentiment, and the diary names a task in two of the nine cells. The skill workshop wrote one general skill in five of the nine cells, none that names a task, and the memory flush never ran, since no conversation was compacted.

The conversation of every earlier instance is kept, but the agent seldom searched it, on 22, 6 and 5 percent of returning instances at low, medium and high effort. At low effort most of its memory searches failed, and that was our configuration's fault. We had left its memory search unset, so it looked for embeddings from a provider the sandbox can't reach, fell back to a keyword index, and then refused later searches against that index. We reran the first 50 instances of three cells with memory search set explicitly. The searches stopped failing, but the agent searched on fewer instances than before, nothing was promoted, and over those 150 instances the budget was unchanged (0.0 points, with a standard error of 2.4). Its tools for reading earlier sessions failed often too, on arguments the model supplied. An earlier session of the same task came back on 7 percent of returning instances at low effort, mostly the instance just before, and those instances cost no less than the memory-off row's.

What the memory-on row does carry is general. The workshop skill was read on 192 of the 600 low-effort stream instances, and it advises the agent to "request another demo when its expected information is worth its cost rather than guessing". The memory-on row buys more examples as a result. At medium effort it buys 0.24 more pairs per instance, on first meetings as much as on returning tasks, which is 1.48 points of its 2.04-point cost. In the low-effort workstreams cell, whose skill appeared at step 17, the memory-on row bought 1.03 more pairs per instance after the skill appeared and none before.

None of the three keeps the examples it paid for

GPT-6 Lunalowmediumhigh
stateless floor3.783.212.77
demo-cache floor2.321.741.26
program library0.590.870.62
prime-agent, persistent3.632.662.53
hermes-agent, persistent3.903.202.14
openclaw, persistent3.572.422.33

Worked examples bought per instance on the tasks that had been met before, in workstreams order. The demo-cache floor is shown its cached pairs at no charge and still buys some. The program library runs its kept program first and rarely needs one. Each harness buys about as many on a returning task as on a new one.

The table shows the one thing all three have in common. None of them keeps the examples it has already paid for, so each buys them again on nearly every returning instance: 2.1 to 3.9 pairs per return on the workstreams stream, no fewer than on a task it has never seen. The demo-cache floor, which is shown its cached pairs for free, still buys 1.3 to 2.3 on a returning task. The three harnesses are built for lessons that carry across unrelated work. None of them ties what it stores to the task id, and none keeps a tested program or the pairs.

Asked to store, two of the three save a few points

Leaving the harnesses alone after the verdict measures them as they come. A reviewer could reasonably object that this sells them short. prime-agent is designed to do well within a session, and nobody told it that knowledge about a task should outlast one. So we ran the invitation to save again, on purpose, as a separate row, and made it as strong as I could write it. After every verdict, the harness now gets this message:

Before you finish, you should liberally save *anything* that *could* potentially help you with future instances of this task, using your own tools. There is no risk in being too verbose: only demonstrations and wrong attempts are scored. Details can always be pruned later if they're surplus to requirements. When you are done, reply with {"action": "finish"}.

prime-agent's version adds the one fact about its store that its rules already carry, that an entry not marked global dies with the session. The rules are otherwise unchanged. A memory-off harness is never asked, since it would have nothing to keep. In all the ablation rows, openclaw runs with its memory search set properly.

who actshow it answersnothing keptkept between instances
bare modelas it choosesstateless floordemo-cache floor: the pairs shown again
bare modela program for every submissionprogram library, wipedprogram library: the pairs, and the kept program run first
harnessas it choosesmemory offmemory on: whatever the harness keeps
harnessas it chooses, asked to storenot runasked to store
harnessthe program library, as a proceduretold to keep programs, memory offtold to keep programs: the programs it saves

The design the ablations fill in, on ARC-AGI-1: who acts, how it answers, and whether anything is kept between instances. Each row on the right is read against its twin on the left. The floors, the program library and the two harness rows of the main results fill four of its cells, and the other rows run only as ablations.

Asked to store, prime-agent saves 3.6, 4.2 and 6.7 points against its memory-off row at low, medium and high effort, where the same harness left alone saves nothing. It saves them on returning tasks, and its first meetings are unchanged within noise. hermes-agent saves 6.4 points at medium effort and 5.2 at high, and nothing measurable at low. openclaw saves nothing at any effort.

Told how to keep a program library, all three save about ten points or more

The invitation says what to save but not how. So we ran one more row, in which the program library's method becomes a procedure that the harness carries out with its own tools. Its rules gain this paragraph:

How to work, in every instance (follow this exactly): keep a program library in the folder {programs}, one Python file per task, named for its task id ({programs}/<task id>.py), defining transform(grid), which takes an input grid as a list of lists of integers and returns the output grid.
1. When an instance starts, look for the task's file. If it exists, run its transform on the test input and submit the result before anything else.
2. If there is no file, or its answer was wrong, request demonstration pairs as you judge necessary, write or fix transform, and run it on every pair you have received until it reproduces all of them; then submit its output for the test input. Checking a program on the pairs is free.
3. Submit only grids that your program produced.
4. After the verdict, save the program that reproduced all the pairs to the task's file, replacing any older version, then finish.

After the verdict, a message also asks it to save the program. We ran the same row with the harness's state wiped before every instance, so that it follows the procedure with nothing to come back to. And we wiped our own program library before every instance, so that it still writes and checks a program for every submission but keeps neither programs nor pairs.

prime-agent, low020406001234567exposure indexprime-agent, medium020406001234567exposure indexprime-agent, high020406001234567exposure indexhermes-agent, low020406001234567exposure indexhermes-agent, medium020406001234567exposure indexhermes-agent, high020406001234567exposure indexopenclaw, low020406001234567exposure indexopenclaw, medium020406001234567exposure indexopenclaw, high020406001234567exposure index
memory offtold to keep programs, memory offmemory on, as it comesasked to storetold to keep programsprogram library

ARC-AGI-1: budget used by exposure index for each harness's ablation rows, one row of panels per harness and one column per reasoning effort, pooled over the three orders, lower is better. The memory-off row (grey, dashed) and the program library (blue) are the anchors, and a row that learns from what it keeps falls from the first towards the second as a task returns. Coral is the harness asked to store and aqua the harness told to keep programs, solid with its memory and dashed with it wiped. The legend runs from the most budget used to the least, though the first three lie within 1.3 points of each other. A point is drawn once at least ten instances fall in it.

Told how, every harness saves more than when asked to store: 9.9 to 13.8 points against its memory-off row at medium and high effort. At low effort prime-agent saves 15.0, openclaw 8.3 and hermes-agent 4.3, the last within noise. At medium and high effort prime-agent's version matches our program library within noise. hermes-agent's and openclaw's fall 2.8 to 4.9 points short, and they cost more than the library on returning tasks, where the library runs its kept program without asking the model anything. At low effort all three fall 11.2 to 19.2 points behind the library, and they answer about half of their returning instances at no cost against the library's two thirds.

asked to store, against memory off-30-20-100+10+20+30points of the budget savedprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhightold to keep programs, against memory off-30-20-100+10+20+30points of the budget savedprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhightold to keep programs, against the same with memory off-30-20-100+10+20+30points of the budget savedprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhightold to keep programs, against the program library-30-20-100+10+20+30points of the budget savedprime-agentlowmediumhighhermes-agentlowmediumhighopenclawlowmediumhigh

What each ablation row saves on the same slots, in points of the budget, pooled over the three orders, with 95 percent intervals from a bootstrap over tasks. Top left: asked to store, against the same harness with its memory off. Top right: told to keep programs, against memory off. Bottom left: told to keep programs, against the same with its memory wiped, which is what keeping the programs adds. Bottom right: told to keep programs, against our program library.

The saving comes from keeping the programs, not from writing them

The wiped variant separates the procedure from the memory. Following the same steps with nothing kept saves at most 3.6 points against the memory-off row, and it costs hermes-agent 8.5 points at low effort. Keeping the programs is what pays: the harness told to keep programs saves 7.0 to 12.8 points against its own wiped twin.

The bare model says the same. Making every submission a program, with nothing kept, saves 17.6, 9.0 and 5.5 points against the stateless floor at low, medium and high effort, because a program checked against the examples catches a wrong rule before it costs an attempt. Keeping the programs and the pairs saves 13.3, 13.7 and 13.4 points more, most of it on returning tasks.

GPT-6 Luna, low020406001234567exposure indexGPT-6 Luna, medium020406001234567exposure indexGPT-6 Luna, high020406001234567exposure index
bare model, nothing keptbare model, pairs kepta program every time, nothing keptprogram library

The bare model's two-by-two on ARC-AGI-1, budget used by exposure index, pooled over the three orders: answering as it chooses (grey) or with a program for every submission (blue), with nothing kept (dashed) or with something kept (solid), its pairs for the demo-cache floor and its pairs and programs for the library.

So kept programs save 7 to 14 points whether our scaffold keeps them or a harness that has been told how. What the harnesses kept of their own accord saved at most 1.0 point, and when they were asked to store, at most 6.7.

openclaw stores when asked, and then never looks

There's a clean way to see whether a learner used what it kept. A returning task answered at cost 0, with no example bought and the first answer right, takes something kept about the task. The memory-off rows almost never manage it, and neither does any harness left to itself.

GPT-6 Lunalowmediumhigh
prime-agent, memory on, as it comes0%0%0%
prime-agent, asked to store13%25%40%
prime-agent, told to keep programs52%72%75%
hermes-agent, memory on, as it comes1%0%0%
hermes-agent, asked to store15%28%41%
hermes-agent, told to keep programs48%67%71%
openclaw, memory on, as it comes0%0%0%
openclaw, asked to store1%3%6%
openclaw, told to keep programs48%65%72%
program library68%84%87%

The share of returning instances on ARC-AGI-1 answered at cost 0, with no example and the first answer right, pooled over the three orders. A learner gets there only with something it kept about the task. The memory-off rows of every harness, not shown, are at 0.2 percent or less.

Asked to store, prime-agent answers 13, 25 and 40 percent of returning instances at cost 0 at low, medium and high effort, and hermes-agent 15, 28 and 41. openclaw manages 1, 3 and 6. Told to keep programs, every harness gets to 48 to 75 percent, and the library to 68 to 87.

openclaw did store when asked. Each of its cells wrote 36 to 74 note files, under names it chose afresh each time. What it lacks is a way back to them. The invitation says what to save and not how to find it again, and each harness decides for itself what comes back into a session. prime-agent's global entries open every session in its digest, and the digest named the returning task on 36, 44 and 63 percent of returning instances. hermes-agent puts its memory and its skill index in every system prompt, and something naming the task reached its model before its first action on 67 to 72 percent. openclaw loads only its MEMORY.md at the start of a session. It put its notes elsewhere and didn't go looking for them: before its first action, it asked for the task on 3.0, 1.3 and 0.2 percent of returning instances. The three cells in which it did write a MEMORY.md hold 47 of its 53 returning instances answered at cost 0.

It's a bit like writing careful notes after every meeting and filing them in a drawer you never open. The program-library rules name the file and say to look there first, and every harness does. It touches the task's program before its first action on 74 to 85 percent of returning instances for prime-agent, 97 to 98 for hermes-agent and 98 to 99 for openclaw. openclaw follows both instructions faithfully, and only the program-library rules tell it where its knowledge will be when it needs it.

Where the notes come back by themselves, asking pays more as the effort rises. Asked to store, prime-agent goes from 13 to 40 percent at cost 0 between low and high effort, and hermes-agent from 15 to 41. openclaw asked for its notes least at high effort.

Nothing transferred to tasks never seen

held-out tasks: budget saved after the stream, against a fresh learner-20-15-10-50+5+10points of the budget (positive means the lifetime helped)program librarylowmediumhighdemo-cache floorlowmediumhighprime-agent, persistentlowmediumhighhermes-agent, persistentlowmediumhighopenclaw, persistentlowmediumhigh

Forward transfer on ARC-AGI-1: the budget each system used on the 25 held-out tasks after its stream, against the fresh reference on the same instances, pooled over the epilogues of the three orders, with 95 percent intervals. Positive means the lifetime made new tasks cheaper. The program library is in blue.

On the 25 held-out tasks, the stream made no system measurably cheaper. The program library saves −0.6, −0.7 and 1.6 points against the fresh reference at low, medium and high effort. Every other interval includes zero except hermes-agent's at low effort, where it uses 9.9 points more after its stream than fresh. That's not a surprise given what the systems keep. The program library keeps one whole program per task, not the primitives the program is built from, and the harnesses keep general advice and a searchable history. ARC's rules share their low-level operations, as the 160 Re-ARC primitives show, and library learning has shown on ARC that abstractions built from solved tasks help with harder ones[18]. So I read the null as a fact about these learners rather than about the benchmark. A learner that refactored what it learned into reusable primitives is the one I'd expect to move this figure.

ARC-AGI-2 and ARC-AGI-3 say the same, with less to go on

GPT-6 Luna, medium020406001234567exposure indexGPT-6 Luna, high020406001234567exposure indexGPT-6 Luna, xhigh020406001234567exposure indexGPT-6 Luna, max020406001234567exposure index
stateless floordemo-cache floorprogram libraryprime-agent, persistentprime-agent, stateless

ARC-AGI-2, workstreams order: budget used by exposure index, one panel per reasoning effort of GPT-6 Luna, lower is better. Instances are the public pairs of evaluation tasks new to the second edition, each visit in a frame the task's rule survives. prime-agent is the only harness on this track, with its memory on (solid) and wiped (dashed); the program library is blue.

On ARC-AGI-2 we ran prime-agent alone among the harnesses, at four reasoning efforts, since a run there costs a good deal more. Paired over the workstreams and sequential orders, its memory saved −0.2, 1.0, −0.1 and 1.5 points at medium, high, xhigh and max effort, all within noise. The floors and the library behave as they do on the first edition. On the workstreams stream the library saves 9 to 13 points against the demo-cache floor at every effort. prime-agent with its memory on uses more of the budget than the demo-cache floor at every effort, and measurably more at xhigh and max.

Claude Opus 5.5, lowstream, budget usedlevels completedheld-out gamessecond meeting of a level
prime-agent, persistent35.138 / 4812.524 → 31 on 15
prime-agent, stateless42.936 / 4813.440 → 32 on 15

ARC-AGI-3, six games over 48 levels with four held-out games, one run of each row on Claude Opus 5.5 at low effort. Budget is actions taken against the human baseline for each level, in percent of a cap of five times the baseline. The last column is the budget on the 15 levels that were met twice, first meeting then second.

ARC-AGI-3 is a different kind of measurement, because a task's next exposure is its next level, which is usually harder, until the game wraps round. With its memory on, prime-agent used 35.1 percent of the budget on the stream and completed 38 of the 48 levels, with seven of its ten failures on one game, sp80. Wiped before every level, it used 42.9 percent and completed 36. Paired level by level that's a saving of 7.8 points with a standard error of 7.7, so within noise for one run of each. The cleanest view of retention is the 15 levels that came round a second time, since a level is the same problem both times. The memory-on run used 31 percent of the budget at the second meeting against 24 at the first, and the wiped run used 32 against 40. So the memory-on run didn't do better at a second meeting than a run that keeps nothing, and there's no sign of a skill kept across a game's levels. With one run each and fifteen levels I wouldn't lean on that reading either way.

What a harness would need

The three harnesses we tested are built for lessons that carry across unrelated jobs. On a stream where the same job keeps coming back, that's the wrong thing to keep. A system that relearns a task on each return pays for the same feedback every time, and the budget records exactly that cost. The program library isn't a sophisticated learner. It keeps one verified, executable thing per task and runs it first, and that was enough to beat every harness by a wide margin.

The ablations narrow down what's missing. Told exactly what to keep and where, the harnesses got most of the way to the library, so their tools aren't what holds them back. My guess at what a harness would gain most from, roughly in order of how cheap each would be to build:

  • Keep the examples it has paid for, keyed on the task id, and show them again when the task returns. The demo-cache floor shows that this alone is worth 7.5 to 10.2 points.
  • Keep a verified, executable rule per task, and run it before thinking about the task at all.
  • Give whatever it keeps a path back into the session, so that the model sees it before its first action rather than only when it thinks to search.
  • Refactor the kept programs into reusable primitives, so that a new task arrives with much of its machinery already in place.

The floors give harness developers a cheap test of any of this. If a mechanism doesn't take a returning task below the demo-cache floor, it hasn't learned the task, whatever else it has stored.

The other opening is the model that drives the harness, and above all what it decides to store. Asked to store, prime-agent and hermes-agent answered about three times as many returning instances at cost 0 at high effort as at low, so a stronger driver already gets more out of the same scaffolding. Frontier models are expensive, though, and plenty of deployments need a model on their own infrastructure. Training a smaller model to drive a custom harness in a custom domain, learning what to store and when to look for it, would bring that ability to those deployments. Continual-ARC is a natural testbed for it, because its domain is one the model largely doesn't know, and its charged verdicts are the reward the training needs.

The benchmark also defines a second regime for exactly this. One system serves a hundred users at once, each on a stream of their own over the same tasks, and learns from all of them between rounds. That configuration yields up to 160,000 verdicts while still scoring every one, and it puts a model trained between rounds on the same scale as a harness that rewrites itself between rounds.

What is still open

The stream is over ARC-AGI-1's public training split, which every developer knows, so it measures generalisation within a task scope that developers can see. The held-out epilogue is the closest we get to the other kind. A request on the epilogue returns one of the task's own curated pairs rather than a generated one, which may be worth more. Both sides of the transfer comparison pay that same price, so I've left it. We don't quantify how hard each task is to generalise, which Chollet lists as ARC's first weakness. ARC-AGI-2 has no generators, so an instance there is a public pair in a new frame rather than a new grid. On ARC-AGI-3 almost no game came round often enough to separate level difficulty from retention.

The stream stands for weeks of work, but each cell ran in under a day, across at most two dates. That compression could have mattered for openclaw's dreaming, which counts recalls on separate days. It couldn't have hidden a promotion, though, because the gate also needs three distinct queries from the agent's own memory searches, and those never came.

Every cell is one run on the fixed stream, with error bars from a bootstrap over tasks, which is also what the leaderboard does. A second seed on any of them would be welcome. The ablations run only on ARC-AGI-1, and the harness asked to store has no memory-off twin, since it would have nothing to keep.

The code, the run records and the paper go out together. The collective configuration, with its hundred users and its 160,000 verdicts, has no runs yet, and that's the next thing.

Numbers from the results mirror at 2026-10-05 10:51 UTC (commit a0ff83f8b); runs complete: ARC-AGI-1 108 of 108, ARC-AGI-2 56 of 56, the ablations 90 of 90.

1. Chollet. On the measure of intelligence. ARXIV 1911.01547, 2019.

2. Karten et al. Prime Agent: a self-improving RLM harness. ARXIV 2608.23552, 2026.

3. Nous Research. Hermes Agent. GITHUB, 2026.

4. OpenClaw contributors. OpenClaw. OPENCLAW.AI, 2026.

5. Akyürek et al. The surprising effectiveness of test-time training for few-shot learning. ARXIV 2411.07279, 2024.

6. Hodel. Addressing the Abstraction and Reasoning Corpus via procedural example generation. ARXIV 2404.07353, 2024.

7. Asawa et al. Continual Learning Bench: evaluating frontier AI systems in real-world stateful environments. ARXIV 2606.05661, 2026.

8. Zhong et al. SkillLearnBench: benchmarking continual learning methods for agent skill generation on real-world tasks. COLM 2026.

9. Shu et al. AgentCL: toward rigorous evaluation of continual learning in language agents. ARXIV 2606.02461, 2026.

10. Joshi et al. SWE-Bench-CL: continual learning for coding agents. ARXIV 2507.00014, 2025.

11. Zheng et al. LifelongAgentBench: evaluating LLM agents as lifelong learners. ARXIV 2505.11942, 2025.

12. Wu et al. StreamBench: towards benchmarking continuous improvement of language agents. NEURIPS 2024.

13. Yang et al. PATH-Bench: path-dependent evaluation of lifelong agents. ARXIV 2608.01149, 2026.

14. Yue et al. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? ARXIV 2504.13837, 2025.

15. Lopez-Paz and Ranzato. Gradient episodic memory for continual learning. NEURIPS 2017.

16. Chollet et al. ARC-AGI-2: a new challenge for frontier AI reasoning systems. ARXIV 2505.11831, 2025.

17. ARC Prize Foundation. ARC-AGI-3: a new challenge for frontier agentic intelligence. ARXIV 2603.24621, 2026.

18. Alford et al. Neural-guided, bidirectional program search for abstraction and reasoning. COMPLEX NETWORKS 2021.