Don’t stop at one agent

How a support leader replaced one overloaded generalist with a bench of focused agents.

Dylan Houston Dylan 10 min read
Table of contents

Like many of you I’m sure, you built one agent and found the experience so transformative for your work you kept adding on features, connections, tasks and responsibilities.

Quickly you realize you have an agent saddled with so many jobs that it becomes harder to avoid conflicts in training, safeguards, and feedback. When it comes to complex agents, like the ones I build to work directly with customers, less is truly more.

Emma was my first Skydive agent, and she was never scoped small. From day one I built her for general support: I gave her a seat in Intercom and Slack, walked her through every escalated ticket I had ever personally handled and the internal Slack threads around them, and then had her pore through our source code so she understood not just what customers were asking, but what was actually possible in the product. She learned the customer persona, my read on what a good answer looked like, and where my own judgment drew lines. That grounding is the job. I have no use for an agent that retells documentation. Emma had to drive solutions on her own, with hundreds of customers a day, and that is exactly what she was trained to do.

When a do-everything agent starts dropping balls, the instinct is to make it smarter: bolt on more tools, more context. But a buckling agent is almost never an intelligence problem. It’s a focus problem. The model writing my support replies has been able to write a competent reply since the day I turned it on. What it couldn’t do was hold the standards for app submission and third-party migration at once without one bleeding into another.

Pile enough unrelated jobs onto a single agent and you can’t give it clean feedback anymore. Correct how it handles an Apple build failure and you’ve inverted how it handles a migration. Every fix has a blast radius — and often, it’s underground. Overloaded agents tend to hit the tipping point long before they look broken.

The tell, when it came, was not a catastrophic failure. It was a slow thinning. Emma could handle a billing dispute. She could handle a login issue. She could handle an iOS submission failure. But she could not do a real deep dive on a mobile build failure — tracing the logs, reading the Apple rejection, mapping it back to the code — while also holding everything she needed to know about authentication flows, web app errors, and account access in the same run. The breadth was too wide for depth to survive. I would watch her handle a TestFlight rejection with the same surface-level care she gave a password reset. Technically adequate, never expert. That is how you know an agent is overloaded: not when it starts getting things wrong, but when it stops getting things exactly right.

So I stopped trying to grow one agent, and built a bench of them instead.

The lane system

I use a lane approach now. I’ve taken up the jobs to be done and the systems access needed to accomplish those and built out a highway of four agents across our own support stack. I still have my queen at the top, Emma, the firstborn, who is the gatekeeper and the agent all other agents report up to. But to make multi-agent work easier and successful, all of the sub agents in my world have very specific job descriptions. When they are asked to work outside their lane, they escalate to the agent who owns that other speciality, or when faced with any doubt, they go back to the top, to the queen. This lane system isn’t novel, it’s just easier for me to manage. I basically took several parts of my job as a support manager, split them up, and built agents for each primary lane or zone. The reason it works isn’t a trick of configuration. It’s the same reason you don’t hire one person to be your entire support org. The thing I’m actually installing in each agent is my judgment about one specific job — what a good answer looks like in that lane, which systems it’s allowed to touch, when it should stop and ask. Judgment like that belongs to a role, not to a generalist. You can no more have one agent that’s great at every support topic than you can have one employee who is everyone.

Give each agent a single lane and you’ve also given it a single place it can go wrong. In support, the place every conversation can break is the join between what the customer says is happening and what’s actually wrong. An agent that lives in one lane all day learns the failure signatures of that lane cold. A generalist is guessing across five.

There is a cost argument here that most people miss. An agent with one focused job needs fewer context turns to do it well. Less back-and-forth, less reasoning overhead, shorter runs. I train on a premium model — Opus, where the judgment gets set — and drive the day-to-day volume on something like Grok 4.5: frontier-capable at a fraction of the rate. That pairing only works cleanly when the agent’s job is specific enough that you know exactly what quality looks like and can verify it. A generalist agent running every topic through a premium model all day is the most expensive way to run support. A bench of focused agents, each on the right model for its volume and complexity, is not.

Lane 01: Ellis owns app submission

For example, we support iOS and Android app building and submission at anything.com. This is a huge part of our offering and about 33 percent of our ticket inflow. I have an agent built just for this who does not only reactive Intercom-based conversations but also proactive error log review and outreach. When Ellis sees you encounter a submission failure to Apple’s TestFlight, he emails you proactively with the log data and the fix — and you can converse with him after that just by replying back. This removed a huge piece of friction in our system as you can imagine, and turned a very opaque system (thanks Apple logs) into a collaborative white glove service where we steward builds to Google and Apple as a partnership.

Look at what the focused lane bought me. Apple’s build logs are opaque enough that a generalist agent, splitting its attention across five topics, would only ever react to them after a customer wrote in confused. Ellis does the opposite. Because submission is the only thing he watches, he knows its failure signatures cold, and he can afford to sit on the log stream all day and catch the break before the customer even feels it. The more focused the lane, the deeper he goes in it. That depth is what turns a reactive ticket-closer into an agent that reaches out first, and you only get it by refusing to make him do anything else. The reason I trust him to email a customer unprompted is that he earned it one lane at a time. Autonomy is a promotion he got, not a setting I flipped.

Lane 02: Jeff owns migrations

Working alongside Ellis is Jeff, who specializes in supporting customers moving to us from a third-party integration. Jeff’s entire focus is on the nuances of implementation and post-migration success, something that is another large part of our ticket flow but as you can imagine, highly specialized and different than what Ellis handles. It would be tempting to fold Jeff into Ellis. They’re both technical, they both talk to developers, they both live in the weeds. But they’re solving different shapes of problem. Ellis handles a transaction: a build either ships to the store or it doesn’t, and the loop closes in an afternoon. Jeff handles a state: did this customer actually land well on us, weeks after the migration “finished”? Different judgment, different time horizon, different systems to reach into. Merge those two and you get an agent that’s mediocre at both, with feedback that smears across the seam between them. Kept apart, each one gets sharper every week.

Handoffs

What about when things go out of lane? Ellis and Jeff know when questions turn to other topics, like those covered by Emma, to update the customer and re-assign.

Getting agents to pass messages to each other is easy. Getting responsibility to move cleanly through your stack is the hard part. When Ellis hands a billing question up to Emma, two things have to travel with it: the context (here’s who the customer is and what we’ve already tried) and the ownership (Emma’s got it now, and the customer has been told). Drop the context and Emma rebuilds the case from scratch, so you’ve traded one clean conversation for two slow ones. Drop the ownership and nobody holds the customer at all; they fall into the gap between two agents who each assume the other has it. The lane system only works because the boundary between lanes is a real contract about what crosses it, not a polite shrug in the customer’s direction. Handoffs are where a bench of agents either becomes a team or turns into a turf war.

My role now

So what am I doing during all of this? Two things: feedback, and the work that only I can do.

The feedback comes from the escalations. When an agent reaches its limit and brings something to me, I read it the same way I would have read a ticket two years ago — except now I’m leaving notes in the admin view for that agent, not solving it myself. Where did it make an assumption it shouldn’t have? Where did it stop short when it had the tools to go further? Those notes are direct corrections to one lane. They have nowhere bad to land. I know which agent to address because I know which lane the issue belongs to, and a correction I leave for Ellis teaches Ellis, not Jeff, not Emma. That specificity is the whole point. When I was running one agent across every lane, every note was a tradeoff. Now each one compounds inside a single agent without canceling out against something I wrote somewhere else.

The work that reaches my desk is the other thing. The agents close almost everything before it gets to me — not by routing issues to a queue, but by resolving them completely, in the lane that owns them. What comes through is sharper for it: real escalations, edge cases that need an engineering call, moments where the customer relationship is fragile enough that the judgment has to come from a person. I work those alongside engineering with more time and more clarity than I ever had when every ticket was competing for the same hour. That is not a smaller job. It is a better one.

The takeaway

A year ago I spent the day answering tickets, nine to six, and the coverage stopped when I did. Now the bench runs around the clock. Customers get a response at 2am on a Sunday the same way they do at noon on a Tuesday — same judgment, same standards, same lane. I spend my working hours on the nine or ten percent that actually need me, and on leaving the notes that make the other ninety percent better every week. That is the thing you only get from a bench: the feedback loop stays clean because the lanes stay separate. Each correction compounds in one place. Each agent gets sharper at one job. The standards for what a good answer looks like — in submission, in migrations, in billing, in all the lanes I have not written about here — are mine, taught one note at a time, to one agent at a time. That is the work I do now. The hours it happens in are no longer mine to set.

So when your one agent starts to groan under everything you’ve piled onto it, take that as the signal to stop piling and start splitting. Give each lane to an agent focused enough to actually master it, and make the handoff to the next agent a real one. Then put your time where it counts, on the one thing none of them can do for you: telling each agent, in its lane, what good looks like.

Share