
AI SEARCH
AI Marketing Agents: How to Build a Six-Agent Team That Gets Your Brand Cited by AI




Written & peer reviewed by Darkroom leardership

AI marketing agents are software workers built on a language model that each own one repeatable marketing job, run on a schedule, and hand their output to the next agent. For AI search, six agents cover the whole loop: find the prompts that matter, audit the answers, fix the pages, earn citations, monitor, and iterate.
Generative AI platforms averaged 9.5 billion monthly web visits between June 2025 and May 2026, up 70% year over year. Every one of those sessions is an answer being assembled, and your brand is either in it or it is not.
Most teams respond by buying a monitoring tool. The tool tells you that you are missing. It does not rewrite the page, pitch the citation or re-test next month, and that is where the work is.
This article expands the six-role roster Darkroom published on social with Blaze Smith of Shovel Studio into something you can build. Darkroom is a growth marketing agency that runs AI search optimization programs for consumer and commerce brands, and the six agents below are the division of labor our own programs follow.
Quick answers
What is an AI marketing agent? A model plus instructions plus tools, given one job and a trigger. It differs from a chatbot in that it acts, on a schedule, without being asked each time.
Do I need six? You need the six jobs done. Start with Scout and Auditor, because everything else depends on their output.
Code or no code? Both work. The Scout, Auditor and Watchtower need a browser or API access to several assistants; the Architect and Editor can live inside a Claude Project or a custom GPT.
What must stay human? Publishing page changes, sending outreach, and deciding what "wrong" means for your brand.
How long before it shows anything? One full cycle to get a baseline, three cycles before movement is readable.
What are AI marketing agents, and how do they differ from an AI tool?
AI agents for marketing are model-driven workers assigned a single job with its own inputs, outputs and trigger. An AI tool answers when you ask. An agent runs because the calendar said so, reads what the last agent produced, and writes something the next agent can use.
That distinction matters more for AI search than for most marketing work, because the job is cyclical. Answers change when models update, when a competitor earns a Reddit thread, and when your own site changes. The unit of work is the cycle, and cycles are what agents are for.
The mistake we see in enterprise AI search programs is one large "GEO assistant" asked to do everything. It produces plausible summaries and nothing to act on. Six narrow agents for marketing, each with one deliverable and one handoff, produce a prompt bank, an audit log, a page brief, a placement list, a trend file and a change list.
If you want the discipline itself defined before the roster, start with what generative engine optimization is. This page assumes that and goes straight to the build.
Why does AI search need a team of agents rather than one assistant?
Because the surface you are optimizing for is fragmented, moving and only partly visible, and a single assistant cannot cover all three at once. A small multi agent system can, provided each agent owns one of those problems.
Fragmented. ChatGPT's share of worldwide generative AI web traffic fell from roughly 76% to 53% over the twelve months to May 2026, while Gemini rose from under 9% to roughly 27% and Claude reached close to 9%. An audit run on one assistant now describes about half the audience.
Moving. In the same Similarweb dataset, the share of US ChatGPT prompts that returned a citation to any website rose from 1.6% to 6.8% over eleven months. The citation layer is being switched on gradually, which means the answer you audited in spring is not the answer being served in autumn.
Partly visible. You only see the answers you test. A brand with no fixed prompt bank is auditing whatever someone remembered to type. That is the gap agentic marketing closes: the agents make the prompt set fixed, the cadence fixed and the logging complete, so the numbers are comparable month to month.
We covered the engine-by-engine trade-offs in our comparative playbook on Perplexity, Claude, Grok and Gemini. The agents below are engine-agnostic by design; the assistant list is a parameter, not a rebuild.
Which six AI agents make up a GEO team?
Six agents cover the AI search loop end to end: Scout, Auditor, Architect, Diplomat, Watchtower and Editor. The table is the roster; the sections after it carry the build notes and a copy-paste prompt for each.
Agent | The one job | Reads | Produces | Runs | Human decision it needs |
|---|---|---|---|---|---|
The Scout | Finds the exact prompts your brand should show up in | Category, product list, competitor list, support tickets, search queries | A ranked prompt bank | Quarterly, and when a new product launches | Approve the category framing and the competitor set |
The Auditor | Checks whether AI even knows your brand exists | The prompt bank | An answer log scored present, missing or wrong, with priority gaps | Monthly baseline, then per Watchtower cycle | Define what "wrong" means for your brand |
The Architect | Builds pages that are easy for bots to grab and quote | Priority gaps, the current page | Answer-first rewrites, FAQ blocks, comparison tables, schema | Per gap | Approve and publish; nothing goes live unreviewed |
The Diplomat | Gets your name onto the sources bots trust most | The Auditor's cited-source list | A placement target list and pitch drafts | Monthly | Send or do not send; relationships stay human |
The Watchtower | Tracks what is working and what is slipping | The prompt bank, prior logs | Mention rate, citation rate and competitor movement over time, with alerts | Fixed schedule, same day each month | Set alert thresholds |
The Editor | Takes what we learn and fixes the plan each time | Watchtower's trend file | A change list routed back to Scout and Architect | Per cycle | Prioritize and reset the clock |
These are ai agent examples in the literal sense: each is a working spec, not a category. The same six types of AI agents apply whether you are a $50M skincare brand or a $500M retailer; what changes is the prompt bank size and how many people sit on the approvals.
Scout: the agent that finds the prompts your brand should appear in
The Scout builds the prompt bank, and the prompt bank is the asset everything else runs on. Its job is to pull the real questions people ask an assistant about your category, test them across ChatGPT, Perplexity, Gemini and Claude, and rank them by how often your brand is missing.
Three inputs make the Scout useful rather than generic: your own data (site search, support tickets, sales-call notes, paid-search queries), because that is how buyers phrase things; the competitor set, because "[brand] vs [competitor]" prompts are where decisions get made; and the assistants' suggested follow-ups, which show how the model frames the category.
The output is ranked by absence: a prompt your brand is missing from on all four assistants outranks one where you appear on three. That ordering is what makes the Scout an AI SEO agent rather than a keyword tool with a new name.
Scout prompt (system instructions). Paste into a Claude Project, a custom GPT, or the instruction field of an agent node.
|

Prompt bank starters. The three patterns from the original post, plus the intents the Scout expands them into.
Pattern | Example prompt | Intent | Why it matters |
|---|---|---|---|
Best [category] brands | "What are the best [category] brands right now?" | Discovery | The shortlist prompt; absence here means you are not a candidate |
Best [product] for [audience] | "What's the best [product] for [audience]?" | Fit | Assistants name fewer brands here, so each slot is worth more |
[brand] vs [competitor] | "[Brand] vs [competitor], which is better?" | Comparison | The decision prompt; the answer often cites a third-party review |
Is [product] worth it | "Is [product] worth the price compared to [competitor]?" | Objection | Where wrong information about you does the most damage |
[use case] without [problem] | "Which [category] works for [use case] without [problem]?" | Fit | Long-tail phrasing that only your support data surfaces |
Auditor: the agent that check whether AI even knows your brand exists
The Auditor runs the prompt bank in full, logs every answer verbatim, and scores each one as present, missing or wrong. Its deliverable is a prioritized gap list, and "wrong" is the category that gets skipped by teams who only count mentions.
Present means the brand is named accurately. Missing means it is not named at all. Wrong means it is named with an incorrect claim: a discontinued product, an old price, a competitor's feature attributed to you, or a stale founder story. A confident wrong answer is worse than absence, because a buyer acts on it.
The Auditor also captures which domains the assistant cited. That list is the Diplomat's input, and it is why the Auditor has to store full answer text rather than a summary. A summary drops the citations.
Scoring is deliberately simple so it stays comparable. Priority is the product of how many assistants the gap appears on and the commercial weight of the prompt, which the Scout already ranked. Comparison prompts with a "wrong" score are P1 by default.
|


Architect: the agent that rewrites pages so bots can grab and quote them
The Architect turns each priority gap into a page change: an answer-first rewrite in plain language, an FAQ block, a comparison table, and schema markup that makes the page legible to a crawler. It drafts; a person publishes.
Answer-first is the whole technique. An assistant lifts self-contained passages, so the first sentence of a section has to be the answer, with conditions after it. Copy that opens with a brand story does not get lifted. The Architect puts the fact first and writes the FAQ in the buyer's phrasing the Scout captured.
Comparison tables are where "wrong" scores get fixed. If an assistant is misquoting your price or a feature, a clearly labeled table on your own page is the correction an assistant can retrieve. The Architect includes scope labels inside cells, so a lifted table cannot misrepresent the data.
Schema is the third output, generated from the page content rather than templated: Article and FAQPage for most pages, Product and Offer for catalog pages.
|

Diplomat: the agent that gets your name onto the sources AI trusts
The Diplomat lists the sites assistants cite most for your category, pitches placements, mentions or reviews on those sites, and tracks which placements actually get picked up in answers. It drafts the pitch; a person sends it.
The source list comes straight from the Auditor's log. If Reddit, a category review site and Wikipedia carry most of the citations behind your competitors' mentions, those are the channels, and no on-site work substitutes for them. The original roster used medium.com, reddit.com and wikipedia.com as the shape; your Auditor data shows the real list.
Two cautions belong in the prompt. Wikipedia and most community sites prohibit brand-driven edits and promotional posting, so the Diplomat targets legitimate contributions and earned reviews, never insertions. And the tracking column is the point: a placement that never appears in an answer is a link, not a citation.
|


Watchtower: the agent that track what is working over time
The Watchtower re-runs the prompt bank on a fixed schedule, logs mention rate and citation rate over time, and raises an alert when a competitor starts winning a prompt you used to hold. It is the AI visibility tracking layer, and it is the agent most teams try to buy rather than build.
Buying is fine for the data collection. The judgment layer is what the agent adds: the same prompt set, the same conditions, the same day each month, and a trend file that separates movement from noise by reporting three-cycle change and labeling anything shorter as provisional.
Four measures cover it: mention rate (the share of prompt-and-assistant pairs naming you), citation rate (the share citing your domain), position in the named list, and the competitor set beside you. What a bad reading means is in AI visibility: how assistants decide which brands to recommend, so the prompt below only specifies the logging.
Alerts are where the Watchtower earns its name. A prompt you held on three assistants last cycle and hold on one this cycle is an alert. A new domain entering the top five cited sources is an alert. A "wrong" score appearing on a prompt that was "present" is the loudest alert of all.
|
Editor: the agent that closes the loop each cycle
The Editor reads the Watchtower's trend file each cycle, sends losing prompts back to the Scout and the Architect, ships the fix, and resets the clock. It is the agent that turns five reports into one operating loop, and it is the one most often skipped, which is why most AI search programs produce dashboards rather than movement.
Its routing rules are simple. Missing with no page behind it goes to the Architect. Present but never cited goes to the Diplomat, because the answer is coming from somewhere else. Drifted phrasing goes back to the Scout. A "wrong" score goes to both the Architect and the Diplomat, because the correction has to land on your page and on the source misquoting you.
The Editor also retires prompts, because a bank that only grows becomes unrunnable: held on all assistants for three cycles with no competitor movement, or irrelevant since the last product change.
|

Which platforms can you build AI marketing agents on?
You can build all six on a no-code AI agent platform, on a model provider's agent framework, or on a mix, and the right choice depends on who will maintain them. The table below covers the platforms we have used or evaluated; each row cites the vendor's own documentation, verified 3 September 2026.
Platform | Type | What it gives the team | Best fit for | Source |
|---|---|---|---|---|
Claude Projects and Cowork (Anthropic) | No code, hosted | A Project holds the system prompt and the fact sheet; Cowork runs scheduled tasks and can drive a browser for the assistant tests. Managed Agents adds hosted, long-running agents via API | Architect, Editor, and a first Scout; Watchtower once scheduling is set | |
Claude Agent SDK (Anthropic) | Code, Python and TypeScript | Subagents, hooks, sessions and MCP connections, so one orchestrator can spawn Scout, Auditor and Watchtower as sub-agents | The full loop as one codebase | |
OpenAI Agents SDK and custom GPTs (OpenAI) | Code, or no code via GPTs | AgentKit launched 6 October 2025 with Agent Builder, Evals, Guardrails and a Connector Registry. Note: OpenAI's documentation now states Agent Builder is deprecated and scheduled to shut down on 30 November 2026, so build on the Agents SDK or custom GPTs, not the visual canvas | Architect and Editor as custom GPTs; orchestration in the SDK | |
Agent Development Kit (Google) | Code, Python, TypeScript, Go, Java, Kotlin | Sequential, loop and parallel multi-agent templates; deploys to Google Cloud or your own containers; works with Gemini natively and other models via adapters | Teams already on Google Cloud who want Gemini in the loop | |
n8n | Low code, self-hostable | An AI Agent node, sub-workflows for multi-agent teams, 500+ integrations, human-in-the-loop checkpoints | Watchtower and Editor, where scheduling and handoffs matter most | |
Zapier Agents | No code, hosted | Agents that run on command or on a schedule across 9,000+ apps, with a knowledge base for the fact sheet | Diplomat tracking, Editor routing into your project tool | |
Gumloop | No code, hosted | Multi-model agent builder with 300+ connectors, scheduled and webhook triggers, web browsing, MCP connections and enterprise controls | Scout and Auditor, where browsing several assistants is the job | |
Lindy | No code, hosted | Slack-native agents with scheduled routines, approval before external actions, 1,000+ integrations and MCP support | Diplomat drafts and Editor summaries delivered where the team already works |
Two practical notes on how to build AI agents for this loop. First, put the brand fact sheet in one place every agent reads, because the Auditor's "wrong" score, the Architect's claims and the Diplomat's proof points all depend on the same source of truth. Second, keep the prompt bank as a versioned file, not a chat history, so a rebuilt agent inherits it.
The platform also matters less than people think, because monitoring tools already do the Watchtower's data collection well; we ranked them in 12 AI search engine optimization tools that deliver results. Build the judgment layer, buy the collection layer, and do not pay twice, which was also the argument in why AEO tools are commoditizing.
How do the six agents run together as one workflow?
They run as a cycle with fixed handoffs, and the AI agent workflow below is the one we operate. The Scout runs first and rarely; the Watchtower runs on the calendar; the Editor runs last and decides what the next cycle contains.
Phase | Agent | Trigger | Handoff |
|---|---|---|---|
Set-up (once, then quarterly) | Scout | New program, new product, or Editor request | Ranked prompt bank to Auditor and Watchtower |
Baseline (once) | Auditor | Prompt bank approved | Scored log and priority gaps to Architect, Diplomat and Editor |
Fix (continuous) | Architect, Diplomat | Editor's change list | Drafts to a human for publish or send; change log back to Editor |
Measure (monthly, same day) | Watchtower | Calendar | Scorecard and alerts to Editor |
Decide (monthly, day after Watchtower) | Editor | Watchtower report | Change list to Scout, Architect and Diplomat; retirements to Scout |
Darkroom runs this cycle inside our AI search optimization service, with the agents orchestrated on Shadow, our AI platform, so the prompt bank, logs and fact sheet stay with the account rather than with whoever ran it last. Why that matters is in what an AI-native agency actually looks like.

What should stay human when you run AI marketing agents?
Four things stay human, and the agents above are written to stop at each of them: publishing a page change, sending an outreach message, defining what a wrong claim is for your brand, and choosing what to fix first.
The reason is not caution for its own sake. Each of the four is where brand, legal or relationship risk sits: a comparison table with an unapproved claim, a pitch that reads as spam to an editor you need next year, a "wrong" definition that flags a true but uncomfortable fact. An agent has no way to weigh those.
Data hygiene is another human job. Every test runs in a fresh session with memory off and a fixed locale, because a personalized answer is not repeatable and an unrepeatable answer cannot be trended. Whoever owns the program owns the conditions.
And a plain limit: we do not yet publish a first-party benchmark for how much mention or citation rate moves per cycle, because we do not have it across enough programs for it to mean anything. When we do, it will carry its prompt count, assistant list, category and period. Treat any vendor quoting a citation lift without those four qualifiers accordingly.
Get the six agents running on your own prompt bank
The fastest route is a baseline: your prompt bank, your category, your competitor set, scored across the four assistants in the first week. That is what our AI visibility audit produces, and the six-agent loop is what runs on it afterward.
Week one. Scout and Auditor: the ranked prompt bank and the scored baseline
Weeks two to four. Architect and Diplomat on the P1 gaps, with your team on approvals
Month two. First Watchtower cycle and first Editor change list
Month three onward. Trend readings, retirements, and the loop running on its own calendar
Darkroom is a growth marketing agency running AI search optimization for consumer and commerce brands, including Cocolab, alongside website and conversion work. See what the first week delivers: AI search optimization at Darkroom.
Frequently Asked Questions
What types of AI agents does a GEO team need?
Six: a Scout to build the prompt bank, an Auditor to score AI answers, an Architect to rewrite pages, a Diplomat to earn third-party citations, a Watchtower to track movement over time, and an Editor to route fixes each cycle. Each has one deliverable and one handoff.
Can you build AI marketing agents without code?
Yes. Claude Projects, custom GPTs, Zapier Agents, Gumloop and Lindy all let you paste the system instructions in this article and attach a fact sheet. Code frameworks such as the Claude Agent SDK or Google's ADK are for teams who want the whole loop orchestrated in one place.
Which agent should you build first?
The Scout, then the Auditor. Every other agent reads their output, and a baseline scored across ChatGPT, Perplexity, Gemini and Claude is what tells you whether the program is working later. Building the Architect first produces content aimed at gaps you have not measured.
How often should the Watchtower re-test?
Monthly, on the same calendar day, under identical conditions: fresh session, memory off, fixed locale, single turn. Re-run within a week of any major page change or model update. Judge movement over three cycles, because two readings cannot separate real change from noise; anything shorter is provisional.
Do AI marketing agents replace an AI visibility tool?
No. Monitoring tools collect answers well and the Watchtower can sit on top of them. What the agents add is the judgment layer: a fixed prompt bank, a scoring rule for wrong answers, routing of each gap to a fix, and a change list a team can act on.
What should never be automated in this workflow?
Publishing page changes, sending outreach, deciding what counts as a wrong claim about your brand, and choosing which gap to fix first. The prompts in this article are written to stop at those points and hand a draft to a person.
Is OpenAI's Agent Builder still a safe place to build?
Build on the OpenAI Agents SDK or custom GPTs instead. OpenAI's documentation states Agent Builder is deprecated and scheduled to shut down on 30 November 2026, with ChatKit remaining available. Anything built on the visual canvas will need exporting before then.

