AEO services are the work of getting a brand named inside AI assistant answers rather than ranked in a list of blue links. A real engagement is five workstreams: recurring visibility measurement, buyer-prompt research, third-party mention and citation earning, content and entity work on your own properties, and technical retrieval hygiene. Anything sold as AEO that produces only a one-off audit or a single visibility score is not a service, it is a screenshot.
"AEO services" gets 320 US searches a month at keyword difficulty 4 with a $51.83 cost per click (DataForSEO, July 2026). Low difficulty, expensive clicks, and a page one made entirely of agencies describing their own offer. Nobody currently ranking has any reason to tell you which parts of the scope are load-bearing and which are filler. That is the gap this page fills.
For the hiring question itself, see how to vet an AI SEO agency. For whether to hire at all, see agency or in-house. This page is only about the scope of work.
The deliverables table
Use this as the spec you hold a proposal against. If a workstream is missing, ask why. If a workstream is present but the verification column cannot be satisfied, it is not being done.
| Workstream | What it produces | How you verify it was done | Could a tool alone do it |
|---|---|---|---|
| Recurring visibility measurement | Share-of-answer per engine, tracked over time on a fixed prompt set, with run-to-run variance shown | Ask for the raw run log: prompt, engine, date, every brand named. Repeated runs on the same date must be visible | Yes, largely |
| Buyer-prompt research | The 20 to 100 questions your buyers actually ask an assistant, mapped to funnel stage | Prompts should look like sentences a customer would type, not keywords with a question mark bolted on | Partly |
| Third-party mentions and citations | New coverage, listicle inclusions, review-site and forum presence, video assets on the sources engines actually cite | Named placements with live URLs and dates, plus the citation share they moved | No |
| Entity and content work | Answer-shaped pages, comparison and alternatives pages, consistent entity descriptions across your properties | Diff the pages. Ask which prompt each page was built to answer | Partly |
| Technical retrieval hygiene | Crawlability for AI user agents, schema, clean extractable markup, no JS-only content | Check robots.txt for AI agents, validate schema, fetch a page with JS disabled | Yes, largely |
| Reporting and interpretation | A monthly read of what moved, what did not, and what changed in the answers themselves | The report should quote actual answer text, not only numbers | No |
The two columns that matter most are the last two together. Where "could a tool alone do it" says yes, you are paying an agency for setup and interpretation, not for labour, and the retainer line for it should be small. Where it says no, that is the actual service.
Which line deserves the biggest share of the budget
Third-party mentions. Ahrefs' analysis of AI visibility found it correlates most strongly with third-party mentions and video, not with on-page work. That is inconvenient for the industry, because on-page work is the part that scales cleanly, bills predictably and can be shown in a deck. Mention earning is slow, partly outside anyone's control, and the least automatable line in the table.
The citation data supports the same conclusion. Reddit is 40.1% of all AI citations, the single most-cited source (Semrush). In Honeyb's own 13 July 2026 measurement, Reddit was 71 of Perplexity's 498 citations and YouTube 40. ChatGPT cited 445 distinct domains across the run, Claude 194, Perplexity 142, and Forbes was the only domain appearing in all four engines' top citation lists. Almost none of that surface is your own website.
So if a proposal puts 60% of the hours into on-page optimisation and schema, and 10% into getting mentioned anywhere else, it is optimising the part that measures easily rather than the part that moves. Ask for the split in hours before you sign.
What is sold as AEO but is standard SEO relabelled
Stated fairly: much of classical SEO does help AI visibility, because retrieval runs on indexed pages. The problem is not that these tasks are worthless, it is that they were already in your SEO retainer and are now being invoiced again under a newer name.
Watch for these:
- FAQ schema as a headline deliverable. Structured data helps parsing. It is not new work and it is not a growth lever on its own. - "Optimised for featured snippets" recast as AEO. Snippet formatting has been standard practice for a decade. - Keyword research relabelled as prompt research. If the deliverable is a volume-sorted keyword list, it is keyword research. Real prompt sets have no volume data because assistants do not publish any. - A one-off audit sold as the engagement. Useful as a starting point, not as a service. See what an AI visibility audit should contain and the GEO audit checklist. - Site speed and Core Web Vitals framed as AI readiness. Worth doing. Not AEO.
The honest version of a proposal says which lines overlap with existing SEO and prices them accordingly. If your current SEO agency is already doing the technical work, you should not be paying twice for it. The distinctions between the disciplines are covered in SEO vs AEO vs GEO.
Want to see this in action?
See how every major AI model talks about your brand. Free to start.
The measurement point everything else rests on
This is the single test that separates a practice from a pitch. The same AI query changes its answer roughly 70% of the time (SparkToro). Honeyb (our product) ran 20 buyer prompts three times each on four engines via API on 13 July 2026, 240 answers in total. The top-ranked brand changed between identical runs on Gemini 44% of the time, Perplexity 43%, ChatGPT 35% and Claude 28%. Brand-set overlap between two runs of the same prompt was Claude 67%, Perplexity 61%, Gemini 54%, ChatGPT 42%.
Top-pick change rate
How often the top recommendation changes between identical runs
An agency that reports a single-run "AI visibility score" is reporting a coin flip and charging you to watch it land. The deliverable has to be a re-measured trend on a fixed prompt set, with the run-to-run variance visible, so you can tell a real movement from the noise floor. If the report shows a number without showing how much that number moves when nothing changes, it cannot be interpreted.
Two follow-on requirements come from the same data. Measurement must be per engine, because the same prompt produced the same top brand on only 20% of ChatGPT and Perplexity pairs versus 53% of Gemini and Claude pairs. And the prompt set must be frozen. Rotating prompts between reports makes every month non-comparable, which conveniently makes every month look like progress.
One more thing worth knowing before you read any report: 62% of AI citations never name the brand being cited (Semrush). Citation counts and brand mention counts are different metrics, and an agency should be able to say which one it is showing you.
What a reasonable engagement looks like in 90 days
| Window | What should be delivered | What should not be promised |
|---|---|---|
| Days 1 to 30 | Frozen prompt set agreed, baseline measured with repeated runs per prompt, technical retrieval issues found and fixed, competitor answer set mapped | Any ranking movement. Thirty days is a baseline, not a result |
| Days 31 to 60 | First trend report against baseline, mention and citation outreach live with named targets, first answer-shaped pages published | Attribution of any traffic change to AEO work this early |
| Days 61 to 90 | Second and third trend reports, first placements landed, share-of-answer movement separated from run variance, budget reallocated toward whatever moved | A guaranteed score. Engines change independently of your work |
The volatility evidence cuts both ways here. Reddit's ChatGPT citation share fell from roughly 60% to roughly 10% in a fortnight in late 2025 (Semrush). A platform-side change of that size can undo or flatter a quarter of work. A good agency says so in advance and shows engine-level data so you can see when a move was theirs and when it was the model's.
What published pricing tells you, and what it does not
Published AEO rates come from agencies' own marketing pages, so treat them as anchors rather than market data. Digital Elevator's pricing guide (last updated 11 May 2026) lists three tiers: $1,000 to $2,500 a month for monitoring and maintenance, $3,000 to $8,000 for active optimisation, and $10,000 to $25,000 and up for enterprise programmes, stating that most credible mid-market retainers land between $2,000 and $10,000 a month.
Read those tiers against the deliverables table. The bottom tier is largely the workstreams a tool can do alone, which is why measurement software sits well below it: Honeyb offers a free check, and monitoring tools range from Otterly at $29 a month to Profound at around $399 a month. The gap between the tool cost and the bottom retainer tier is what you are paying for interpretation, so ask what the interpretation consists of.
Before you brief anyone
Get your own baseline first, so the agency's opening measurement is something you can check rather than something you have to accept. Decide whether the answer is even a priority for your category by reading is answer engine optimization worth it, and get the mechanics from how to get cited by AI.
Then run the check yourself. Honeyb (our product) will measure how often AI assistants name your brand across engines, free, at /tools/ai-visibility-checker. Bring the result to the first call and ask the agency to reproduce it.





