Forty agents was never the plan. The plan was to stop being the bottleneck.
Part of our guide to AI agents for small business.
WE•DO ran the way most small agencies run. A handful of people. More than thirty active accounts. A month-end that ate a week.
Client reports took five hours each. A full SEO audit took a day. A content brief took ninety minutes before anyone wrote a sentence. Every one of those jobs was necessary, repeatable, and completely unstrategic.
By the end, you'll know:
- What nine of our named agents do, and the five traits they all share
- How the four-layer stack fits together (and why the model is the cheapest part)
- How much time the fleet gave back on four recurring jobs
- The four times it went wrong, including the one that still stings
- How to pick your first agent with four questions
It's a 13-minute read. If you only want the first-agent checklist, skip to the last section.
The obvious fix was to hire. We ran that math and didn't like it. Building the capability in house costs $105,500 to $167,500 a year. That counts salary, benefits, tools, training, and management overhead. It also takes six to twelve months to reach productivity. All that buys one person doing the same mechanical work slightly faster.
So we wrote a job description instead. Then another. Today the fleet runs 40+ agents across content, publishing, SEO, reporting, meeting prep, and internal operations. It's how a team this size delivers 100+ client reports a month with no reporting department.
01 / why-build-agents-instead-of-hiring
Why Build Agents Instead of Hiring?
Because our bottleneck was hours, not ideas. Nobody at WE•DO was short on strategy. We were short on the time strategy needs. Those hours went into data aggregation, formatting, and follow-up.
Three constraints pushed the decision:
- Cost. A hire absorbs the work at full loaded cost. Cost scales linearly, and nothing compounds.
- Single-person risk. When one person holds a process in their head, it leaves when they do.
- The work that never happened. Weekly performance checks, pre-call prep, unlogged time. Real value, always last in line.
The reframe that made it work: an agent is an employee and a skill is an SOP (standard operating procedure). A new hire is only as good as the procedure you hand them. So is an agent.
That vocabulary is the spine of the delegation framework we run on. It also kept us from building what everyone builds first: one do-everything assistant. It's mediocre at five jobs and trusted for none.
02 / what-do-40-agents-look-like
What Do 40+ Agents Look Like?
Four functions, one shared shape. Nine named agents are below. That's enough to show the pattern without turning this into a directory.

Content and publishing
Content Cal runs the blog pipeline end to end. That covers keyword research, SERP analysis (a look at what already ranks), brief, draft, QA, and the review handoff. It reads Search Console, GA4, and DataForSEO before it writes a word. It stops at review every time. This post is one of its drafts.
Publishing Pat owns deployment. It commits the post and its graphics, opens the pull request, gets a verified preview URL, and waits. It can't publish until a human says yes in the task thread. That rule exists because of a failure further down this page.
SEO and research
SEO Sage handles audits and keyword scoring. It pulls a crawl and cross-references Search Console. It scores opportunities on volume, difficulty, and commercial intent. Then it writes the survivors into tasks with the evidence attached.
The full-day audit now takes 30 minutes. Writing speed isn't why. Five data sources arrive at once.
A competitor watch agent runs alongside it on a schedule. It tracks gaps and SERP movement so nobody has to remember to look.
Reporting and analytics
Reporting Roger is the highest-ROI agent we've built, and the least interesting. It pulls GA4, Search Console, and Ads, then calculates period-over-period change. It drafts the analysis and renders it into one of 43 branded report templates. A person reviews it, adds the context only they have, and ships it.
An anomaly agent watches for metric shifts between reports. A 15% traffic drop reaches us before it reaches a client on a call.
Client operations
Intake Ike routes inbound work into the right list with the right fields filled in.
Briefing Blair starts 24 hours before any client call. It pulls the last three meeting notes, open tasks, and analytics changes. Then it posts a prep brief to the task.
Billable Time Sweep finds unlogged hours before the week closes. It's the only agent here that pays for itself in revenue, not time.
Different jobs, identical shape. Every agent in the fleet has:
- One job, described in one sentence with no "and" in it.
- One trigger that isn't a person remembering. A status change, a due date, a schedule, an inbound form.
- Two or three trusted sources, not ten. Extra context means extra chances to reason from something irrelevant.
- One output format, so "good" is checkable.
- One owner who feels the pain and complains when it stops.
We killed the agents missing two of those five. The pattern held often enough that we wrote it up in why most AI agent pilots fail.
03 / how-is-the-stack-built
How Is the Stack Built?
Four layers, and the model is the only one we swap. Most "how to build an AI agent" guides start with the model. Ours starts with the floor. The model is the cheapest thing in the stack to change. The platform is the most expensive.

The work platform is the floor
Our tasks, docs, comments, approvals, and history live in one system. That one decision moves output quality more than any prompt we've written. An agent can reach real context without a scavenger hunt. Spread that context across five tools and you pay for the sprawl on every run.
It also means the work lands where the work already happens. An agent that posts a draft into a chat window creates a copy-paste tax. An agent that comments on the task doesn't.
Connectors are what turn a writer into an operator
Without integrations (connections to your other tools), an agent is a chatbot with a longer prompt. With them, it can read the account and act on it.
The fleet reaches GA4, Search Console, DataForSEO, meeting transcripts, the site repository, and the blog publisher. The full inventory and wiring is in how we use Claude Code inside our operations.
This is the line between assisted and autonomous. Can it retrieve and act, or does it wait for you to paste? We draw that same line in AI tools for business automation.
Skills are the SOPs
A skill is a written procedure any agent can load. Ours cover brand voice, how a post moves from research to QA, deliverable formatting, and graphics. Refine a procedure once and every agent that loads it improves at the same time.
This is the part that compounds. It's also the part teams skip. Prompts don't accumulate. Documented procedures do.
The model is a setting
Nobody knows which model will be best in six months, or what it will cost. So the model is a dropdown per agent, not a foundation. The employee stays the same. So do the SOP and the memory. Only the engine changes.
04 / what-does-a-typical-week-look-like
What Does a Typical Week Look Like?
Monday starts without anyone opening a tool. That's the most useful thing about the fleet, more than any single output.

Every lane ends at a human. Reports are drafts until someone signs off. Posts sit in review until someone approves. Client emails stay unsent. Delegation without oversight is negligence with extra steps.
05 / how-much-time-did-the-fleet-give-back
How Much Time Did the Fleet Give Back?
A monthly client report went from 5 hours to 30 minutes. The other gains look the same: repeated work, nothing clever. That's boring, which is how you know it's real.

| Job | By hand | With the fleet | What changed |
|---|---|---|---|
| Monthly client report | 5 hours | 30 minutes | Data pulls and first-draft analysis |
| Full SEO audit | 1 day | 30 minutes | Five sources arriving together |
| 12-site WordPress audit | 2 hours | 15 minutes | One command across every site |
| Content brief | 90 minutes | 10 minutes | Research and structure pre-assembled |
One usage snapshot from a heavy stretch: about 12,000 platform AI credits used. In the same period, the fleet returned 63 hours of work. Put a loaded rate on those hours and the argument stops being theoretical.
The four-line cost breakdown and break-even model are in what running 40+ agents costs and saves. So are the cases where it goes negative.
The condition
Recovered hours only count if they move to work that earns
63 hours returned that turn into 63 hours of cleanup is not a saving. The gain is real when the time lands on strategy, pitches, delivery, and experiments.
06 / what-didnt-work
What Didn't Work?
Four things broke, and each one produced a written rule. Every post about an agent fleet shows the wins. Here are our documented failures.

The draft that skipped the research
One agent wrote a full draft before the keyword and SERP work was done. The prose was fine. The angle targeted a keyword with ten searches a month. We killed that version and labeled it so nobody would reuse it. Research is now a gate the draft can't skip.
The pull request that pretended to be a preview
Our publishing agent returned a GitHub pull request link and called the post ready for review. A pull request isn't a preview. The reviewer got a code diff, and sometimes a 404.
The fix was a written rule. No handoff ships without a verified, login-free preview URL. A pull request link may never be described as one.
The post that shipped with no task behind it
A post went live with no task in the pipeline. So it had no owner, no publish date on record, and no performance check scheduled. We caught it days later and created the task after the fact. Now every publish creates its own tracking record before it can be marked done.
29,085 impressions and 8 clicks
This is the one that stings. Over the last 180 days, our 40 AI marketing agents pillar page pulled 29,085 impressions and 8 clicks in Search Console. That's a 0.03% CTR (click-through rate) at an average position of 10.6.
Our top 25 blog pages in the same window: 78,813 impressions, 77 clicks, and 617 pageviews in GA4.
The query report makes it worse. A meaningful share of those impressions come from junk long-tail strings, not buyers.
The fleet did exactly what we asked. It published more, faster, more consistently. It couldn't fix our aim: broad topics with no commercial intent.
Agents multiply whatever strategy you aim them at. Aim them at the wrong thing and you get a bigger wrong thing, on schedule.
07 / why-was-agent-30-easier-than-agent-3
Why Was Agent 30 Easier Than Agent 3?
Because by agent thirty, the scaffolding already existed. The third agent took weeks. The thirtieth took an afternoon. The models didn't change that much in between.
A new agent now inherits the connectors, the brand voice skill, the QA standard, and the approval behavior. It also inherits a shared definition of a finished deliverable. Building one is mostly writing a job description and naming a trigger.
The real asset is the standards under the agents. If a fire took the whole fleet tomorrow, the documented procedures would rebuild it in a week. Lose the procedures and 40 agents become 40 unmaintainable one-offs.
08 / where-should-a-small-business-start
Where Should a Small Business Start?
Start with one agent, attached to a job that already happens every week. You don't need forty. Forty is what three years of a team compounding on its own standards looks like.
Pick it with four questions:
- Does this job happen at least weekly? Frequency pays back the build.
- Is it context-heavy? Docs, tasks, transcripts, analytics. Context is where agents beat tools.
- Is "good" easy to define? A structured output can be checked. A vague one can't.
- Does a missed step cost something? That's where reliability turns into money.
Three yes answers means build a narrow first version. Give it a trigger. Connect only the sources it can't work without. Name the number it should move. Then hand it to whoever feels the pain today.
That's the whole method. It isn't glamorous. It's the only version we've seen survive a real workweek.
Next step
Find the first job worth handing off
We will map your recurring workflows, price the two or three worth automating first, and tell you plainly which ones to leave alone. Book an AI workflow audit at wedoworldwide.com/services/ai-operations.