5 Mar 2026

Demonstration without a mode.

Teaching an agent by showing it is usually built as a mode, with a capture pipeline and a store of learned skills that belong to it alone. We argue that it should fall out of three general properties instead, namely images that are ordinary message content, task sessions that stay open for corrections, and a library that any running task can write to.

Teaching by showing is usually a mode

The obvious way to let a person teach an agent by showing it is to build a mode for it. The person presses a button to start, a dedicated pipeline captures the screen and the narration, a model turns the capture into a skill, and the skill lands in a store of things learned by demonstration. We first built something close to this, a component that captured keyframes aligned with the person's narration and produced a transcript with its images, which the component started and stopped around the demonstration. We removed it, and teaching by demonstration now uses no component of its own.

A mode has three limitations

Firstly, a mode keeps its own copy of state and storage. What it captures and what it learns live in structures of its own, which the rest of the agent never searches or calls. Secondly, a recording is finished when it stops. A correction such as "no, not that button, the one underneath" then becomes an edit to a finished artefact, and not a message to someone still doing the work. Finally, every mode multiplies the combinations the agent has to handle. A screen shared during a recording, or a scheduled task that fires in the middle of a demonstration, each needs an answer of its own.

a modepress start, and later stopcapture keyframes with the narrationa model writes the skilla store of demonstrated skillsthree general propertiesimages asmessagessessions thatstay openstorage fromany taskthe libraryfunctions and guidanceevery other task

Two ways to learn from a demonstration. A mode, on the left, has a capture step, a model and a store that belong to it alone, in coral. On the right, three general properties, images as messages, open sessions and storage from any task, lead into the same library as everything else, in blue.

Images are ordinary content

The property that does most of the work sits well upstream of teaching. A shared screen is not recorded. The latest frame is kept, at most one per second, and converted to an image only when it is needed. When the person speaks, the frame is taken and joins the next state the conversation model reasons over, paired with what the person said at that moment. The moment the person speaks is usually the moment the picture is worth having, and the frames in between are discarded.

A frame from a webcam takes the same path, and so does a screenshot of the agent's own desktop. Each becomes an image part in a message, the same kind of thing as a sentence, and none of these paths knows what the image will be used for.

Pixels travel by reference

Each message that came with frames carries an annotation naming the files they were saved to. When the conversation forks into a task, the image data is removed from the history it passes on, and the annotations survive. The conversation model is told that the task can reach an image only through its path, and to include every relevant path in its request, from earlier messages as well as the latest one.

The task then chooses which frames to look at. It opens a frame in its sandbox and displays it, which puts the image into its own context in the same way as any image its code produces. Most tasks open none of the frames, and a task learning a visual workflow may open nearly all of them. Twenty frames passed along with every fork would make every prompt large, and paths let each task pay only for the frames it needs.

the conversation“Click Export, here.”the pixels[Screenshots: frames/0412.jpg]the fork passes the pathand drops the pixelsthe task…frames/0412.jpg…display(Image.open(path))opened onlywhen needed

A frame passed by reference. The conversation holds the image and the path it was saved to. The fork passes the path, in blue, and drops the pixels, and the task opens the frame from its path only when it needs to see it.

A teaching session is an ordinary session

Suppose the person shares a screen and says "this is how I do the weekly invoice run". The conversation model sees the words and the frames, decides that work is needed, and keeps the task open, because more instructions from the person are likely. Its instructions list walkthroughs as one example of such work, and no instruction on the path treats a demonstration differently from any other request. The task opens the frames it needs and works out the procedure. When the person says "no, not that button, the one underneath", the correction is an interjection into a session that never ended, and the task still has the whole context to apply it.

A session that never ends never reaches the review that follows a finished task. A task can instead ask for a review of its trajectory so far at any point, naming the part it wants kept. The review writes to the same two libraries as the review after any task, a repeatable sequence becoming a function and a rule becoming guidance linked to it. The person can keep correcting, and the task can update the function in place instead of writing a second one.

the personconversationthe taskthe library“This is how I do the invoice run.”“No, not that button, the one underneath.”frames, each with a pathstarts a task, kept openinterjectsopens frames 2 and 5, works it outapplies itstore_skillsthe session stays open throughouta function, with guidance linked

A teaching session in ordinary parts. The person talks through the invoice run with frames attached, the task opens the frames it needs, a correction arrives as an interjection, and the task asks for its trajectory to be stored, in blue, without ever finishing.

What the session produces

The result is an ordinary function, with any guidance linked beside it. Nothing marks it as learned by demonstration, it sits in no separate library, and it runs by no separate path. It can be called by another function, found by a search when something related comes up, or rewritten later by a task that never saw the screen. A skill learned through a pipeline of its own would have lived in that pipeline's store, where none of these could reach it.

Cypher's collection on programming by demonstration gathers systems that infer programs from a user's actions[1]. Argall et al. survey robots that learn policies from demonstrations, and the choices each makes about how a demonstration is recorded[2]. Tesler argued against modes in interactive software, since a mode changes what the same action means[3]. Raskin traces a class of user errors to modes that the user has lost track of[4]. Of the work we know, the systems that learn from demonstration each give it a dedicated channel for recording. We argue that an agent can do without one when its images, its sessions and its library are general enough to carry a demonstration between them.

Open questions

Firstly, frames are taken when the person speaks. A click made in silence leaves no frame unless the person says something while its result is on screen, and a demonstration given without narration is mostly invisible. Secondly, the conversation model has to copy every relevant path into its request, and a path it leaves out is a frame the task can never see. Finally, nothing checks that a function learned from a demonstration reproduces the demonstration. The task judges its own understanding, and the first real test of the function is its first real use.

However, we argue that these concern what the general parts capture and check, and that none of them calls for a mode. We leave those improvements to future work.

1. Cypher (ed.). Watch what I do: programming by demonstration. MIT Press, 1993.

2. Argall et al. A survey of robot learning from demonstration. Robotics and Autonomous Systems, 2009.

3. Tesler. The Smalltalk environment. BYTE, 1981.

4. Raskin. The humane interface. Addison-Wesley, 2000.