Training an agent with a Loom

A screen recording can transfer the intent, actions, and judgment that a written procedure leaves out.

Dylan Houston Dylan 6 min read
Table of contents

There’s a moment in training a new agent where writing it all down stops being worth it. I can feel it coming. I’m about to teach an agent something that lives in my hands more than my head, a task with a dozen little decisions I make without narrating them to myself, and I can already tell the written spec is going to be a page long, take me forty minutes, and still miss the three things that actually matter. That’s the moment I stop typing and hit record.

I’ll give you a small one. I needed to hand off a task to Georgia, one of my agents, that came down to a connection check: log into a thing, run it, confirm it actually worked rather than just claiming it saved. I could have written that out. “Navigate to the dashboard, authenticate, trigger the test, verify the result.” But which screen? Which button? What does “verify” mean here, the green checkmark or the actual round-trip? Every one of those gaps is a place Georgia would have had to guess, and a guess in a connection check is exactly the kind of confident-but-wrong I train hardest against.

So instead I recorded seventy-seven seconds of myself doing it, talking the whole way through. “Test, log in, verify.” And that short clip carried more usable signal than the page I would have written, because I got to show and tell at the same time, in the exact order the task happens, without stopping to spell out every step.

Why the video beats the spec

Words alone leave gaps. When I write “check the connection,” the agent is filling in blanks I didn’t know I left. A Loom closes them, because it hands over two things at once and lets the agent line them up against each other.

The narration carries my intent. When I say “log in and verify” out loud, I’m telling Georgia why she’s clicking, not just where. The screen carries the ground truth. My cursor moving through the task removes the ambiguity that written steps always leave, because now there’s a right answer to “which button,” and it’s the one I clicked. And the agent cross-checks the two. It hears the reason, sees the action, and where they line up there’s nothing left to interpret. That’s the part that makes it click. I’m not describing the task from the outside, I’m doing it, and Georgia is watching the doing.

It’s also just faster and more honest for me. Some tasks I understand with my hands, not in prose, and forcing them into a written procedure is where I introduce the errors, because I skip the step that’s obvious to me and invisible to everyone else. Recording myself doing it for real keeps the obvious steps in, including the small decisions I’d never have thought to write.

And the reason this works now, when it wouldn’t have a couple of years ago, is that the newest models actually watch. Georgia doesn’t just read a transcript of what I said. She takes the video in frame by frame and listens to the audio alongside it, so she can see my cursor land on a button at the same second I say why I’m clicking it. That’s the cross-check: she lines up what she heard against what she saw, frame to word, and where they agree she knows she understood it right. A written spec gives her one channel to trust. A Loom gives her two that confirm each other.

What actually transfers

Here’s the part that surprised me the first time. From that one short Loom, Georgia didn’t just pick up the steps. She picked up the judgment, which parts of that check genuinely needed me and which she could take off my plate entirely. That’s the thing you can’t get into a checklist. A list of clicks tells an agent what to do; it doesn’t tell it which moments are load-bearing and which are rote. When I narrate a task while I run it, the emphasis comes through, the “this is the part that actually matters” lands, and the agent comes away knowing where the risk lives.

That’s the same thing I’m always after when I train an agent by any method: not a retelling of the procedure, but the reasoning behind it, so it can handle the case I didn’t record. A Loom just happens to be the fastest way I’ve found to hand that over. I show how I think about the task, not only what I click, in the time it takes to record it once.

How I record one that holds

The method is short, because the whole point is that it’s cheaper than writing.

Do the task for real. Not a staged demo, the actual thing, start to finish, including the small forks where I decide something. The real sequence is the signal. A cleaned-up demo drops exactly the judgment I’m trying to transfer.

Narrate the intent, not just the motions. “I’m logging in to confirm the connection really works, not just that it saved” teaches more than ten seconds of silent clicking. The why is the part the screen can’t show, so I say it out loud.

Call out where a human belongs. Somewhere in most tasks there’s a seam, a moment where I’d want the agent to stop and check with me instead of pushing on. I name it while I’m there. “If this comes back ambiguous, don’t guess, flag me.” That one sentence is the difference between an agent that finishes everything and one that knows when finishing is the wrong move.

Hand it over and ask for the writeup. I share the Loom and tell the agent to turn what it saw into a repeatable process it can run the same way next time, saved to its own operating notes. Now the seventy-seven seconds isn’t a one-off, it’s a lane of expertise the agent owns.

The takeaway

A Loom lets me hand an agent how I work, not just a list of instructions. I show and tell at once, the agent cross-checks the two against each other, and it comes away with the reasoning behind the task instead of a brittle set of steps. For anything complex enough that the written spec would be long and still incomplete, which is most of what’s worth delegating, recording myself is the faster, clearer way to train. Georgia learned that connection check from a clip shorter than this paragraph took you to read, and she’s run it ever since without me in the loop.

And here’s the part that changes the math on your whole library: you’ve probably already made videos like this. Every onboarding walkthrough, every training clip, every screen recording you cut for a new hire is now training material for your agents too. The same video that showed a person how the job is done shows an agent the same thing, watched frame by frame and cross-checked against your narration. You don’t have to start from scratch. Point your agents at what you already recorded.

Share