7 May 2025

Storing skills as functions.

An agent that works over a long period meets the same kinds of task again and again. We argue that what it learns about doing them should be kept as a library of Python functions rather than a list of notes. Functions call one another, so shared steps and shared abstractions each exist once, and every task that needs them calls them.

An agent meets the same task many times

An agent that works for someone over a long period will be asked to do the same kinds of task many times, each time with small differences. The second time it searches a website for posts on a topic should cost less than the first, and the tenth should cost very little. Whether it does depends on what the agent kept from the earlier attempts, and in what form it kept it.

We think an agent's memory should be split by what is being remembered, since each kind of memory has a structure of its own. What was said suits a transcript of every exchange, with summaries over it. Facts suit tables whose columns can be added to and restructured as new facts arrive. What is to be done suits a list of tasks, with their deadlines and schedules. The fourth kind is how to do things, which a planner draws on each time it plans a task. Its form is the least settled of the four.

the agentplans each task and carries it outreads and writeswhat was saidtranscripts and summariesfactstables it can restructurewhat to do, and whentasks and their scheduleshow to do ita library of functions

An agent's memory, split by what is remembered. Each store keeps its contents in the form that suits them. The fourth, in blue, keeps how to do things as a library of functions, which the agent reads and writes as it plans a task.

A list of notes is flat

Most memory for language model agents is kept as text. MemGPT moves text between the context window and an external store, in the way an operating system pages memory[1]. Reflexion writes reflections on failed attempts and reads them before the next attempt[2]. Agent Workflow Memory is closer to a store of procedures, as it induces reusable workflows from past trajectories on the web, but it also keeps them as text for the model to read[3]. Text suits a fact, such as the year someone was born. For a procedure, a list of notes has three limitations.

Firstly, a list of notes is flat. Each note describes one whole procedure, and a step that several procedures share is written out again inside each of them. Nothing in the store records that the steps are the same, and an improvement to one copy does not reach the others. Secondly, notes do not build on one another. In code, a pattern that recurs is factored out into a function with a name, and the function is called wherever the pattern is needed. That function can itself be built from smaller ones, and over many such refactorings the code settles into shared primitives at the bottom and shared abstractions above them. A note can mention another note, but it cannot run it, and the model still has to assemble the whole procedure from the pieces every time. Finally, a note cannot be checked on its own. Whether it is correct shows only when a model reads it and acts on it, and a failed run does not say whether the note was wrong or the model misread it.

Functions form a hierarchy

We propose to store each skill as a Python function in a shared library. Its signature and docstring say what it does, and its body says how, by calling other functions. At the bottom of the call stack, each function can be a plain English command such as "click the search bar". A controller then maps each command onto a single action in the environment, choosing among the actions that its current state allows. In a browser, that action is a single click or key press. The figure below shows three skills kept both ways.

As notes
Posts about a topic: go to linkedin.com, click the search bar, type the topic and press enter, then click the Posts filter.

People with a name: go to linkedin.com, click the search bar, type the name and press enter, then click the People filter.

Messaging someone: go to linkedin.com, click the search bar, type their name and press enter, click the People filter, then click Message on the first result, type the message and press enter.
As functions
def search_linkedin(query: str):
    """Search LinkedIn for `query`, starting from the homepage."""
    go_to_linkedin_com()
    click_the_search_bar()
    type_the_text(query)
    press_enter()


def send_message(text: str):
    """Message the first person in the results with `text`."""
    click_message()
    type_the_text(text)
    press_enter()


def find_posts_about(topic: str):
    """Show the LinkedIn posts about `topic`."""
    search_linkedin(topic)
    click_the_filter("Posts")


def find_people_named(name: str):
    """Show the LinkedIn members called `name`."""
    search_linkedin(name)
    click_the_filter("People")


def message_person(name: str, text: str):
    """Send `text` to the LinkedIn member called `name`."""
    find_people_named(name)
    send_message(text)

Three skills kept as notes (top) and as functions (bottom). The notes are an illustration of what a text memory might keep. Every note repeats the search, which the functions write once as search_linkedin. message_person is built from another task, find_people_named, and the lowest-level calls, such as click_the_search_bar(), are English commands for the controller.

A function calls many functions, and many functions call it. This many-to-many structure is loosely analogous to Wikipedia, where a page links to many other pages and is linked to from many more. Each topic is written once, on its own page, and every page that needs it points to it. A library of functions is organised the same way. search_linkedin is written once and called by every task that searches, and find_people_named is a task in its own right as well as a step inside message_person. The analogy breaks in one respect. A link on Wikipedia is a reference that a reader may follow, while a call is part of the procedure itself. A change to a function changes every task that calls it, for better or for worse.

message_person(name, text)find_posts_about(topic)find_people_named(name)send_message(text)click the filtersearch_linkedin(query)go to linkedin.comclick the search bartype the textpress enterclick Message
a stored functionan English command for the controllermore than one caller

The functions above as a call graph, with an arrow from each function to what it calls. Coral marks every function or command with more than one caller, and the coral arrows are the calls into them. A repair to one of these reaches all of its callers, and deleting it removes them too.

We expect this structure to matter more as the library grows. The same low-level steps recur across the tasks on one website, and similar patterns recur across websites. When each is a function, a new task can be written mostly as calls to functions that already exist, and a repair to a shared step reaches every task above it. Wang et al. compared programs and text as the form of web agents' skills on the WebArena benchmark[4]. Their agents with skills as programs did better than the same agents with skills as text by 11.3% in success rate, and took 10.7 to 15.3% fewer steps than their baselines. They attribute most of the gain in success to verifying each skill by running it, and the saving in steps to composing primitive actions, such as a click, into higher-level skills, such as searching for a product.

A function can also be checked, which a note cannot. It either raises an exception or finishes, and a finished function can be checked against its docstring. In a hierarchy, the check can run at every level, and a fault can be found in the function where it happens instead of only in the task at the top.

A library needs its call graph

The library should record the hierarchy as data. Each function can be kept as one record that holds, next to its source, the name of every function it calls. Before a function is stored, its source can be parsed and rejected unless four rules hold.

  • The source contains exactly one top-level function.
  • The function imports nothing.
  • It calls nothing through an attribute, such as math.sin() or posts.append().
  • Every name it calls is one of 30 permitted builtins, such as len and sorted, or another function supplied with it.

We ask for these rules so that a stored function can only do what the library and the permitted builtins let it do. Everything it calls then appears in its stored list of calls, and the call graph of the library is complete. Deletion depends on this. Removing a function should also remove every function that calls it, directly or further up. The library then never holds a function that would fail on a missing call. The same graph tells us which tasks a repair will reach.

A planner should not have to read every function before it can use the library. A listing can return the signature and docstring of each function, and the source only when asked for it. A search can filter the library with an expression over any stored field. A filter such as 'search_linkedin' in calls returns every function that calls search_linkedin, which is the library's equivalent of the list of pages that link to a page on Wikipedia. The planner can then read the whole catalogue cheaply, and read the source only of a function it is about to call or change.

Writing functions from the top down

A library of this kind can be grown one task at a time, from the top down. Given a function's signature and docstring, a model writes its body. It is told to call functions that do not exist yet whenever it lacks the context to finish, or a modular solution would be clearer, and to give each of them an expressive name. Each missing function is filled in when the program first reaches it. The call raises a NameError, and the traceback gives every caller above it and the line where each makes its call. From that context, the model proposes a docstring and a signature for the new function, which is then implemented in the same way. Code as Policies writes programs for robot control in this way, recursively defining each function that a program calls but does not define[5]. Parsel likewise decomposes a task into a hierarchy of described functions, implements each with a code model and tests them together[6]. This continues down the call stack until the functions at the bottom are simple browser actions, and a person can review each function as it is written.

Checking every level

A plan can itself be Python source, written using the functions already stored. When the description of a task matches the docstring of a stored function exactly, the stored function can stand in for the whole plan. Every function in the plan, at every level, is then wrapped in a check that it did what its docstring says. A function that raises an exception is re-implemented without a check. A function that finishes is checked by a model, which reads everything that ran below it, down to the browser actions and their screenshots, and decides whether the task in the docstring was done. When the check fails, the function can be re-implemented, or the check can step up and rewrite its caller if the structure above is wrong. The task is complete once the check passes for the function at the top.

run the functionraisedre-implement itrun againfinishedcheck it against its docstringa model reads everything that ran below it,down to the browser actions and screenshotsfailedfailedrewrite the callerwhen the structure aboveit is wrongpassedreturn to the callerat the top of the call stack,the task is complete

The check around every function in a plan. Blue marks the path to a completed task, and coral the paths taken after a failure. A failed check either re-implements the function or rewrites its caller, and the rewritten function is run and checked again.

Voyager keeps a library of executable JavaScript skills for Minecraft and composes new skills from old ones[7]. It adds a skill only after a self-verification step judges that the skill completed its task, and it retrieves skills by the embedding of their descriptions. Of the work we know, it is the closest to the argument here. LATM has one model write tools as Python functions for another model to use[8]. TroVE induces a toolbox of verified functions for programmatic tasks, which it grows as it uses them and trims periodically[9]. CodeAct has agents act by writing Python directly, instead of emitting fixed tool calls[10]. DreamCoder learns a library of reusable abstractions by refactoring the programs it finds[11]. Agent skill induction, discussed above, brings verified programmatic skills to web agents[4]. SkillWeaver has web agents explore a new website, practise the skills they discover and distil them into a growing library of APIs[12]. A-MEM gives a text memory some of the structure we want, by linking each new note to related notes in the manner of a Zettelkasten[13]. Its links record that two notes are related, while a call records that one procedure is made of another.

In each of the works with programs, the lowest level is a fixed set of operations, such as Mineflayer calls in Voyager or clicks on element ids in WebArena. Functions that bottom out in English commands are instead resolved by the controller against the environment as it is when the command runs. We expect this to let a function survive small changes to a page, such as a renamed button, but we have not measured it.

Open questions

The hierarchy does not arise by itself. Writing one task from the top down gives one hierarchy per task. Turning these into shared abstractions needs a refactoring step, which looks across the stored functions for a pattern that several of them repeat, factors it out, and rewrites the callers to use it. DreamCoder and TroVE each address a version of this problem[11][9], and we do not know which approach suits a library of agent skills.

The check depends on a model judging screenshots against a docstring, and we do not know how often it is wrong. A false pass costs more in a library than in a note, because every task that calls the function inherits the fault. Retrieval is also open. An exact match between a task and a stored docstring would miss a request worded differently, and some form of search by meaning will be needed. Deleting every caller of a deleted function keeps the library runnable, but it also removes callers that could have been repaired, and marking them for re-implementation may be better. Banning attribute calls makes some code awkward, since posts.append(post) has to be written as posts = posts + [post], and we accept this because it keeps every call visible to the check.

However, we argue that none of these questions weighs against storing skills as functions. Each concerns how the library is searched and maintained, and we leave them to future work.

1. Packer et al. MemGPT: towards LLMs as operating systems. ARXIV 2310.08560, 2023.

2. Shinn et al. Reflexion: language agents with verbal reinforcement learning. NEURIPS 2023.

3. Wang et al. Agent workflow memory. ARXIV 2409.07429, 2024.

4. Wang et al. Inducing programmatic skills for agentic tasks. ARXIV 2504.06821, 2025.

5. Liang et al. Code as policies: language model programs for embodied control. ICRA 2023.

6. Zelikman et al. Parsel: algorithmic reasoning with language models by composing decompositions. NEURIPS 2023.

7. Wang et al. Voyager: an open-ended embodied agent with large language models. ARXIV 2305.16291, 2023.

8. Cai et al. Large language models as tool makers. ICLR 2024.

9. Wang et al. TroVE: inducing verifiable and efficient toolboxes for solving programmatic tasks. ICML 2024.

10. Wang et al. Executable code actions elicit better LLM agents. ICML 2024.

11. Ellis et al. DreamCoder: growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning. ARXIV 2006.08381, 2020.

12. Zheng et al. SkillWeaver: web agents can self-improve by discovering and honing skills. ARXIV 2504.07079, 2025.

13. Xu et al. A-MEM: agentic memory for LLM agents. ARXIV 2502.12110, 2025.