How to choose an agency for visibility in AI assistants
Illustration: AI-generated
A while ago our own Search Console showed a query no human types into Google. A German B2B software vendor described in full sentences that their audience increasingly researches in ChatGPT rather than Google, that they want to be recommended there, and that they lack the in-house knowledge for it. Then they asked which agencies in Germany specialise in visibility inside AI assistants and also handle ongoing work, and explicitly asked for concrete names.
That was not a person at a keyboard. That was a language model searching the web on that vendor’s behalf. This is what demand looks like now: the research happens inside the assistant, and whoever is not named there is not on the list.
Plenty of agencies have responded by putting "GEO", "AI visibility" or "visibility in AI assistants" on their website. A good share of them still do classic SEO and hope it somehow carries over to language models. This article shows you how to spot real specialisation, which questions to ask in the first call, and what even the best agency cannot promise.
What separates GEO from classic SEO
Classic SEO optimises for rankings and clicks. Generative engine optimisation optimises for an AI naming, citing or recommending your brand inside a synthesised answer. That is not a difference of degree.
| Aspect | Classic SEO | Real GEO work |
|---|---|---|
| Goal | Positions 1 to 3 on Google | A mention or citation inside the AI answer |
| Success metric | Rankings, traffic, click-through rate | Citation rate, share of voice, recommendation frequency |
| Key signals | Backlinks, keywords, technical SEO | Entity clarity, authoritative third-party sources, structured facts |
| Measurement | Search Console, rank tracking | Prompt-based testing across several models |
| Time horizon | 3 to 6 months | Weeks on-page, months for off-page authority |
An agency that talks about "AI SEO" and hands you ranking reports has not yet understood the difference. [1][2]
The key finding: authority beats technology
The largest cross-platform study on this so far measured 2,089 brands across ChatGPT, Claude, Gemini and Perplexity. Its central result: brand authority predicts visibility in AI answers about three times more strongly than technical optimisation of your own website. Expressed as correlations, 0.42 against 0.14; in explained variance, 17.6 percent against 1.9 percent. [4]
The clearest illustration is a pattern the study calls the "GEO paradox": companies with excellent technical setup but weak brand authority stay effectively invisible. They have schema markup, clean meta tags, an llms.txt, and still average a visibility score of 14 out of 100. The reverse also holds: brands with strong authority and mediocre technology get recommended reliably.
An independent analysis of 75,000 brands reaches the same conclusion and names the strongest individual signals: mentions on YouTube correlate most strongly with visibility at around 0.74, branded web mentions at around 0.66. Both sit outside your own website. [6]
A note on the evidence, because this article would otherwise do exactly what it warns against: the 2,089-brand study comes from a vendor that sells a visibility product. Its sample is self-selected, overrepresents tech and SaaS companies, and deliberately includes unfiltered personal blogs, weekend projects and freshly registered domains. Its widely quoted headline, that 85.7 percent of all brands are invisible, is therefore overstated. The authority finding still holds, because an independent study of 75,000 brands points the same way.
1. A real GEO methodology, not repackaged SEO
A specialised agency can explain how it structures content so that language models extract and cite it. It talks about entity optimisation, schema stacking, source gap analysis and prompt clusters. If the first call only covers keyword density and link building, you are getting SEO with a new label. [1]
2. Measurement across several engines
It tracks at least ChatGPT, Perplexity, Gemini or Claude and Google AI Overviews, at prompt level and on a schedule. Whether that runs on specialised tooling or an in-house dashboard matters less. What matters is that a defined prompt set is queried repeatedly, rather than someone occasionally taking a look. [3]
3. Demonstrable change
It shows you before-and-after examples with named prompts and documents what exactly was changed. "We are seeing great results" is not an answer. Check that the evidence shows rates across many runs rather than single screenshots, because one answer is chance. [3]
4. A plan for authority beyond your website
Given the data, this is the single most important point. AI answers lean heavily on third-party sources: trade media, directories, review platforms, YouTube, community discussions. A good agency has a concrete plan for getting you into those, not just a content roadmap for your own blog. [4][6]
5. Technical retrievability as the baseline
Schema.org implemented properly, an unambiguous entity definition, crawler accessibility, llms.txt where useful. This is the entry ticket, not the competitive edge: its absence can disqualify you, its presence alone does not set you apart. Anyone selling you technology as the main deliverable is selling you the 1.9 percent. [1][5]
6. Transparent metrics
Before the engagement starts it is clear what gets measured: citation rate as the share of relevant prompts in which you are named, share of voice against defined competitors, the position of the mention, and where possible downstream effects such as referral traffic from AI assistants. [1]
7. Honest expectation setting
It tells you what is not possible. Guaranteed mentions or "number one in ChatGPT within 30 days" are not ambitious targets. They signal that someone either does not understand the mechanics or is counting on you not understanding them. [3]
Eight questions for the first call
- Show me a client where you measurably increased citation rate. What exactly did you change?
- What monitoring do you use for ChatGPT, Perplexity, Gemini and Claude, and how often do you run it?
- How do you decide which prompts get prioritised?
- How do you handle the fact that the same question returns different answers on different days?
- Which levers outside our website do you use to build authority signals?
- What does your process for schema and entities look like?
- What do you measure after 30, 60 and 90 days, and what does success look like concretely?
- Can you run five real buying questions from my audience right now and show who gets named today?
The last question is the most effective one. Anyone who cannot do that on the spot does not work with this daily. [3][5]
What cannot be controlled
Most providers leave this part out. It belongs in any honest basis for a decision.
- Individual answers are not predictable. In a study with roughly 600 participants and about 3,000 runs, the chance of getting the same brand list twice was lower than 1 in 100. The same list in the same order was closer to 1 in 1,000. What is stable is the frequency across many runs. That is precisely why serious work measures rates, not positions. [7]
- Without third-party sources, little happens. Your own website alone is rarely enough. If trade media, directories and communities do not mention you, the models stay reserved. [4][6]
- Existing prominence matters. An audit across roughly 37,000 runs found that brands in the lowest prominence tiers never surfaced at all in 48 to 52 percent of cases. Starting small does not mean starting at zero, it means starting below it. [8]
- Every engine behaves differently. In the 2,089-brand study, citation rates for ChatGPT and Gemini correlated at only 0.19, ChatGPT and Claude at 0.49. A brand named often in one assistant faces close to a coin flip in the next. A single strategy for all models fails on this. [4]
- Authority takes time. Entity building and mentions beyond your own site work over months. Technical improvements land faster but do not replace reputation.
- The framing is not yours. Even when you are named, the model decides in what context and in which words. Outdated pricing, long-retired features or plainly false claims about you are not the exception but the norm. [4]
- Measurement stays incomplete. There is no Search Console equivalent for AI assistants. Every tool works from prompt samples, including the large ones: the broadest public index evaluates 126 million prompts and is still a sample. [9]
Anyone selling you a fixed position inside an AI answer is selling hope, not craft.
How to proceed
- Write down 20 to 40 real buying and comparison questions from your audience. Not questions about your brand name, every model answers those.
- Run those questions yourself today in ChatGPT, Perplexity and Gemini. Note who gets named. That is your baseline and it costs you an afternoon.
- Build a shortlist of three to four agencies and ask them the eight questions above.
- Weight measurement and methodology highest in your decision, not presentation and price.
- Start with a clearly scoped 90 day pilot with a documented baseline. Without one, nobody can prove progress to you later.
A short note on us
This article is written as a neutral selection guide, which is why the self-promotion sits in one place rather than spread through the text.
Grovia Digital is a performance marketing and AI consulting agency based in Dubai, focused on the DACH and MENA markets. We work with B2B companies on getting named and recommended in ChatGPT, Perplexity and other assistants, through measurable citation work, entity optimisation and ongoing support. Whether that fits your situation is worth a conversation. You are welcome to put the eight questions above to us exactly as you would to any other agency.
Key Takeaways
- 1Brand authority predicts visibility in AI answers about three times more strongly than technical optimisation of your website. Technology is the entry ticket, not the edge.
- 2Assess an agency against seven criteria and weight measurement and methodology highest. The most effective question: have them run five real buying questions live.
- 3Individual AI answers are not predictable, frequency across many runs is. Measuring positions instead of rates means measuring noise.
- 4Guaranteed mentions do not exist. Start with a 90 day pilot and a documented baseline.
Sources and context: the figures on brand authority, the GEO paradox and platform divergence come from the Loamly study (2,089 brands, February 2026). It is vendor research, its sample is self-selected and unfiltered; we deliberately do not quote its headline that 85.7 percent of brands are invisible, because personal blogs and test domains inflate it. The core authority finding is supported by an independent Ahrefs analysis of 75,000 brands. All correlations are associations, not demonstrated causes.