Some of what an agent learns cannot be run
We argued earlier that the procedures an agent learns should be stored as functions, which call one another and can be checked. Some of what an agent learns is not a procedure, though. A rule about the tone of messages to investors, a preference for the staging database in experiments, or a walkthrough of an application's settings screen does not execute. It still has to be stored where the agent will find it when it applies, next to the code it governs.
The common answer is the skill folder. In Anthropic's Agent Skills, a skill is a directory holding a document with a name, a description and instructions, and optionally scripts beside it[1]. The agent sees every skill's name and description at the start, reads the full document when a task matches, and reaches the scripts through the document. Loading the details only when they are needed is the right instinct, and the folder is easy to package and share. We found that the shape of the folder, one document with code nested inside it, sets three limits on what the library can express.
A folder of prose has three limitations
Firstly, a rule that applies to several skills has no single home. Suppose a tone rule governs both drafting emails and building decks. It is either pasted into both skills, where the copies drift apart, or kept as a skill of its own, which the agent has to remember to load alongside the task. The format has no way to say that one rule governs several capabilities. Secondly, code is found only through the document around it. A script whose purpose matches the task is missed when the description of its skill does not, and skills that wrap a single script need a document that exists only to make the program retrievable. Finally, the folder is a tree. Each script belongs to exactly one skill, where the relation between rules and code is naturally many-to-many.
The same knowledge kept as skill folders, on the left, and as two linked libraries, on the right. In the folders the tone rule is copied into two skills, in coral, and drifts. In the libraries it is one guidance entry, in blue, linked to the three functions it governs, and each function can be found on its own.
Two libraries, linked both ways
We propose to keep two libraries. A function is an entry of its own, with a name, a signature, a docstring and the implementation, and it is found by searching for its meaning. A piece of guidance is an entry of its own too, holding procedural knowledge, such as operating procedures and walkthroughs of software, and it is searched in the same way. Its text can carry images aligned with it, such as screenshots of the settings screen that a walkthrough describes, which code cannot hold.
The link between the two is explicit and many-to-many. Each piece of guidance lists the functions it concerns, and each function lists the guidance that refers to it. Both lists are foreign keys, and deleting a function removes it from the guidance that referred to it, as deleting guidance removes it from its functions.
{
"guidance_id": 7,
"title": "Writing to investors",
"content": "Lead with the number that changed. Keep to one screen. "
"Never forecast beyond the current quarter.",
"function_ids": [12, 31, 47], # draft_email, build_deck, post_update
}
The tone rule then becomes one entry linked to the three functions it governs, and editing it once changes what all three are told. Nothing forces a pairing. Guidance with no functions, such as the preference for the staging database, attaches to nothing executable, and most functions carry no guidance at all, since a good docstring already says what they do.
Both libraries are searched first, as peers
Retrieval treats the two libraries as equals. At the start of a task the agent may call nothing but the two searches, and the rest of its tools appear only once it has searched both. A function found by a search is placed straight into the agent's sandbox with the functions it depends on, and the agent can call it at once, with no document between the search and the code. Guidance found by the same step arrives with the ids of its functions, and a function arrives with the ids of its guidance. Either one leads to the other.
The first steps of a task. Until the agent has searched both libraries, in blue, those searches are the only tools it has. Once it has, every other tool appears, and the function it found can be called directly.
We made the search compulsory because an agent that is free to start work at once can rebuild from scratch a procedure that the library already holds. Requiring both searches, and not either one, keeps the agent from finding the code and missing the rule that governs it.
The review writes to both
After each task, a review reads what happened and decides what, if anything, should join either library. Its instructions describe the two stores as the what and the how. A function is a single callable, like a tool with its docstring, and guidance is a recipe for a multi-step workflow, like a prompt that refers to tools. Guidance is written only when a composition would be hard to rediscover, and never to repeat what a docstring says. When a trajectory yields both, the review stores the function first, then guidance that points at it by id.
What the review can do with a finished task. Most often it stores nothing. It may store a function, write guidance with no code, or store a function and then write guidance that links to it, in blue.
The order matters because the link is made by id. Guidance written before its function exists has nothing to point at, and a function stored after its guidance is not linked to it unless the guidance is updated.
Related work
Agent Skills package a skill as a folder of instructions with optional scripts, loaded progressively as a task needs them[1]. Voyager's skills are JavaScript programs for Minecraft, retrieved by the meaning of their descriptions[2]. ExpeL extracts insights in natural language from an agent's successes and failures and recalls them in later tasks[3]. Agent Workflow Memory induces reusable workflows, written as text, from past trajectories[4]. CoALA distinguishes procedural memory, held in an agent's code and in the model's weights, from semantic memory, which holds knowledge about the world[5]. The systems we know keep what an agent learns in one form, as code or as text, or nest code inside text. We keep the two as separate libraries of equal standing, with links between them that either side can follow.
Open questions
Firstly, the review decides what is a function and what is guidance, and it can divide a lesson badly. A rule that should have been a branch in the code may end up as prose that the code never reads. Secondly, the links are kept on deletion but not on change. A function rewritten so that a rule no longer applies to it still carries the link, and nothing checks that the guidance still fits. Finally, folders share better than our libraries do. A directory can be zipped and installed anywhere, and registries have been built on that, while our libraries live in a database and are shared by scope.
However, we argue that these are reasons to maintain the links with care, and not to nest code inside prose again. Each concerns how the two libraries are kept consistent and shared, which we leave to future work.
1. Anthropic. Equipping agents for the real world with Agent Skills. Anthropic engineering blog, 2025.
2. Wang et al. Voyager: an open-ended embodied agent with large language models. ARXIV 2305.16291, 2023.
3. Zhao et al. ExpeL: LLM agents are experiential learners. AAAI 2024.
4. Wang et al. Agent workflow memory. ARXIV 2409.07429, 2024.
5. Sumers et al. Cognitive architectures for language agents. TMLR 2024.