How We Built 40+ AI Agents to Run Our Agency
Strategy

How We Built 40+ AI Agents to Run Our Agency

The roster, architecture, weekly cadence and four failures behind the 40+ AI agents that run WE•DO. First-party numbers, no demos.

Forty agents was never the plan. The plan was to stop being the bottleneck.

Part of our guide to AI agents for small business.

WE•DO ran the way most small agencies run. A handful of people, more than thirty active accounts, and a month-end that ate a week. Client reports took five hours each. A full SEO audit took a day. A content brief took ninety minutes before anyone wrote a sentence. Every one of those jobs was necessary, repeatable, and completely unstrategic.

The obvious fix was to hire. We ran that math and did not like it. Building the capability in house costs $105,500 to $167,500 a year once you count salary, benefits, tools, training, and management overhead, and it takes six to twelve months to reach productivity. That buys one person doing the same mechanical work slightly faster.

So we wrote a job description instead. Then another. Today the fleet runs 40+ agents across content, publishing, SEO, reporting, meeting prep, and internal operations, and it is how a team this size delivers 100+ client reports a month with no reporting department.

Here is the whole thing: what each agent does, how the stack is put together, what a week actually looks like, what it gave back in hours, and the four times it went wrong. Including the one that still stings.

Why We Built Agents Instead of Hiring

The bottleneck was never creative. Nobody at WE•DO was short on strategy. We were short on the hours strategy needs, because those hours were going into data aggregation, formatting, and follow-up.

Three constraints pushed the decision:

  • Cost. A hire absorbs the work at full loaded cost and scales linearly. Nothing compounds.
  • Single-person risk. When one person owns a process in their head, the process leaves when they do.
  • The work that never happened at all. Weekly performance checks, pre-call prep, unlogged time. Real value, always last in line.

The reframe that made it work: an agent is an employee and a skill is an SOP. A new hire is only as good as the procedure you hand them, and so is an agent. That vocabulary is the spine of the delegation framework we run on, and it kept us from building the thing everyone builds first: one do-everything assistant that is mediocre at five jobs and trusted for none.

The Roster: What 40+ Agents Actually Look Like

The fleet splits into four functions. Nine named agents are below, which is enough to show the pattern without turning this into a directory.

Four function groups with nine named WE•DO agents: content and publishing, SEO and research, reporting and analytics, and client operations.

Content and publishing

Content Cal runs the blog pipeline end to end: keyword research, SERP analysis, brief, draft, QA, and the review handoff. It reads Search Console, GA4, and DataForSEO before it writes a word, and it stops at review every time. This post is one of its drafts.

Publishing Pat owns deployment: it commits the post and its graphics, opens the pull request, gets a verified preview URL, and waits. It cannot publish without a human saying yes in the task thread. That constraint exists because of a failure further down this page.

SEO and research

SEO Sage handles audits and keyword scoring. It pulls a crawl, cross-references Search Console, scores opportunities on volume, difficulty, and commercial intent, then writes the survivors into tasks with the evidence attached. The audit that used to be a full day is 30 minutes, and the difference is not writing speed, it is that five data sources arrive at once.

A competitor watch agent runs alongside it on a schedule, tracking gap and SERP movement so nobody has to remember to look.

Reporting and analytics

Reporting Roger is the highest-ROI agent we have ever built, and the least interesting. It pulls GA4, Search Console, and Ads, calculates period-over-period change, drafts the analysis, and renders it into one of 43 branded report templates. A person reviews, adds the context only they have, and ships it.

An anomaly agent watches for metric shifts between reports, so a 15% traffic drop reaches us before it reaches a client on a call.

Client operations

Intake Ike routes inbound work into the right list with the right fields filled in. Briefing Blair starts 24 hours before any client call, pulls the last three meeting notes, open tasks, and what changed in the analytics, and posts a prep brief to the task. Billable Time Sweep finds unlogged hours before the week closes, which is the only agent on this list that pays for itself in revenue rather than time.

Different jobs, identical shape. Every agent in the fleet has:

  • One job, describable in a sentence with no "and" in it.
  • One trigger that is not a person remembering. A status change, a due date, a schedule, an inbound form.
  • Two or three trusted sources, not ten. Extra context is extra chances to reason from something irrelevant.
  • One output format, so "good" is checkable.
  • One owner who feels the pain and complains when it stops.

Agents missing two of those five are the ones we killed. That pattern is consistent enough that we wrote it up separately in why most AI agent pilots fail.

The Architecture: Four Layers, One Swappable Model

Most "how to build an AI agent" explanations start with the model. Ours starts with the floor, because the model is the cheapest thing in the stack to change and the platform is the most expensive.

Four-layer AI agent architecture stacked from skills to agents to connectors to the work platform.

The work platform is the floor

Our tasks, docs, comments, approvals, and history live in one system. That single decision moves output quality more than any prompt we have ever written, because an agent can reach real context without a scavenger hunt. Spread that context across five tools and you pay for the sprawl on every run.

It also means the work lands where the work already happens. An agent that posts a draft into a chat window creates a copy-paste tax. An agent that comments on the task does not.

Connectors are what turn a writer into an operator

Without integrations, an agent is a chatbot with a longer prompt. With them, it can read the account and act on it. The fleet reaches GA4, Search Console, DataForSEO, meeting transcripts, the site repository, and the blog publisher. The full inventory and how it was wired is in how we use Claude Code inside our operations.

The line between assisted and autonomous is exactly here: can the thing retrieve and act, or does it wait for you to paste? That distinction is the one we draw in AI tools for business automation.

Skills are the SOPs

A skill is a written procedure any agent can load: how our brand voice works, how a blog post moves from research to QA, how a client deliverable gets formatted, how a graphic gets built. Refine the procedure once and every agent that loads it gets better at the same time.

This is the part that compounds, and it is also the part teams skip. Prompts do not accumulate. Documented procedures do.

The model is a setting

Nobody knows which model will be best in six months, or what it will cost. So the model is a dropdown per agent, not a foundation. The employee stays the same, the SOP stays the same, the memory stays the same. Only the engine changes.

What a Typical Week Looks Like

The most useful thing about the fleet is not any single output. It is that Monday starts without anyone opening a tool.

Five weekdays showing which WE•DO agent fires each day and what it produces.

Every lane ends at a human. Reports are drafts until someone signs off. Posts sit in review until someone approves. Client emails stay unsent. Delegation without oversight is not delegation, it is negligence with extra steps.

The Numbers: What the Fleet Actually Gave Back

The honest version of the ROI story is boring, which is how you know it is real. The gains came from repeated work, not from anything clever.

Four recurring jobs compared by hand and with the agent fleet: monthly report five hours to thirty minutes, SEO audit one day to thirty minutes, twelve-site WordPress audit two hours to fifteen minutes, content brief ninety minutes to ten minutes.

JobBy handWith the fleetWhat changed
Monthly client report5 hours30 minutesData pulls and first-draft analysis
Full SEO audit1 day30 minutesFive sources arriving together
12-site WordPress audit2 hours15 minutesOne command across every site
Content brief90 minutes10 minutesResearch and structure pre-assembled

One usage snapshot from a heavy stretch: roughly 12,000 platform AI credits consumed against 63 hours of work returned in the same period. Put a loaded rate on those hours and the argument stops being theoretical. The full four-line cost breakdown, the break-even model, and the cases where it goes negative are in what running 40+ agents actually costs and saves.

The condition

Recovered hours only count if they move to work that earns

63 hours returned that turn into 63 hours of cleanup is not a saving. The gain is real when the time lands on strategy, pitches, delivery, and experiments.

What Did Not Work

Every post about an agent fleet shows the wins. Here are four documented failures from ours, and the rule each one produced.

Four documented WE•DO agent failures and the written rule each one produced.

The draft that skipped the research

One agent wrote a full draft before the keyword and SERP work was done. The prose was fine. The angle was aimed at a keyword with ten searches a month. We killed the version, labeled it so nobody would reuse it, and made research a gate the draft cannot skip.

The pull request that pretended to be a preview

Our publishing agent returned a GitHub pull request link and reported the post as ready for review. A pull request is not a preview. The reviewer got a code diff, and in some cases a 404. The fix was a written rule: no handoff ships without a verified, login-free preview URL, and a pull request link may never be described as one.

The post that shipped with no task behind it

A post went live with no task in the pipeline, so it had no owner, no publish date on record, and no performance check scheduled. We caught it days later and created the task retroactively. Now every publish creates its own tracking record before it can be marked done.

29,085 impressions and 8 clicks

This is the one that stings. Over the last 180 days our 40 AI marketing agents pillar page pulled 29,085 impressions and 8 clicks in Search Console, a 0.03% CTR at an average position of 10.6. Across our top 25 blog pages in the same window: 78,813 impressions, 77 clicks, and 617 pageviews in GA4.

Read the query report and it gets worse. A meaningful share of those impressions come from junk long-tail strings, not buyers. The fleet did exactly what we asked. It published more, faster, more consistently. It could not fix the fact that we pointed it at broad topics with no commercial intent.

Agents multiply whatever strategy you aim them at. Aim them at the wrong thing and you get a bigger version of the wrong thing, on schedule.

The Compound Effect: Why Agent 30 Was Easier Than Agent 3

The third agent took weeks. The thirtieth took an afternoon. Nothing about the models changed that much in between. What changed is that by agent thirty, the scaffolding already existed.

A new agent now inherits the connectors, the brand voice skill, the QA standard, the approval behavior, and a shared definition of what a finished deliverable looks like. Building one is mostly writing a job description and naming a trigger.

The asset is not the agents. It is the standards underneath them. If a fire took the whole fleet tomorrow, the documented procedures would rebuild it in a week. Lose the procedures and 40 agents become 40 unmaintainable one-offs.

What This Means If You Run a Small Business

You do not need forty. Forty is what happens after three years of a team compounding on its own standards. You need one, attached to a job that already happens every week.

Pick it with four questions:

  • Does this job happen at least weekly? Frequency is what pays back the build.
  • Is it context-heavy? Docs, tasks, transcripts, analytics. Context is where agents beat tools.
  • Is "good" easy to define? A structured output can be checked. A vague one cannot.
  • Does a missed step cost something? That is where reliability turns into money.

Three yes answers means build a narrow first version. Then give it a trigger, connect only the sources it cannot work without, name the number it should move, and hand it to whoever feels the pain today.

That is the whole method. It is not glamorous, and it is the only version we have seen survive contact with a real workweek.

Next step

Find the first job worth handing off

We will map your recurring workflows, price the two or three worth automating first, and tell you plainly which ones to leave alone. Book an AI workflow audit at wedoworldwide.com/services/ai-integration.

Book an AI workflow audit.

Frequently Asked Questions

How many AI agents does a small business actually need?
One, to start. A single agent attached to a weekly, context-heavy job with a definable output will prove or disprove the model faster than a fleet. Most teams that try to launch five at once ship none of them reliably.

What is the difference between an AI agent and an AI tool?
A tool waits for you to open it and paste context. An agent has a trigger that is not a person remembering, access to the data it needs, and an output format someone can check. If you cannot name what starts it without a human in the sentence, it is a tool.

Which AI agents should a marketing team build first?
Reporting and content briefs, in that order. Both repeat on a known cadence, pull from several sources, and are easy to start manually and easy to never finish. That combination is where the payback shows up fastest.

Do you need engineers to run 40 AI agents?
No, but you need documented procedures. The fleet is maintained by marketers, because the hard part is writing the SOP and defining the trigger, not writing code. What you cannot skip is someone owning instruction updates and QA.

What does it cost to run a fleet of AI agents?
Runtime is the smallest line. Build and maintenance carry the weight, and platform choice moves ROI more than model choice. The four cost lines and the break-even math are broken down in our AI agent ROI piece.

Do AI agents publish work without a human?
Not in our operation. Every agent output is a draft until a person approves it. Reports stay private links, posts stay in review, and client emails stay unsent. That gate is the reason the fleet is trusted with real accounts.

About the Author
Mike McKearin

Mike McKearin

Founder, WE-DO

Mike founded WE-DO to help ambitious brands grow smarter through AI-powered marketing. With 15+ years in digital marketing and a passion for automation, he's on a mission to help teams do more with less.

Want to discuss your growth challenges?

Schedule a Call

Continue Reading