Generative engine optimization: getting cited, and knowing whether it worked
Every agency in this category will tell you they can improve your share of voice in ChatGPT. Almost none will tell you that no answer engine exposes a ranking API, that every tool in the industry samples rather than observes, that roughly 30% of brands stay visible from one run of the same prompt to the next, or that LLM referral traffic is currently around 0.1% of website traffic. All four are true, all four are on this page before any sales material, and the work is still worth doing — for the right brand, at the right stage, measured honestly.
The short version
- What it is: generative engine optimization — making a brand findable, parseable and citable by ChatGPT, Perplexity, Google AI Overviews and the other answer engines. Also sold as AEO, LLM SEO and AI visibility; the labels differ more than the work does.
- What it costs: quoted per programme, because the work genuinely differs by category and corpus footprint. Our tracking product, measure.co, is published at $89/month and stands alone.
- What nobody can promise: placement. There is no ranking API and no supplier controls what a model says. Anyone guaranteeing citation is selling you something they cannot deliver.
- What the honest measurement looks like: every tool in this category samples rather than observes, LLM referral traffic is currently around 0.1% of website traffic, and roughly 30% of brands stay visible from one run of the same prompt to the next. We lead with that rather than bury it.
- Who should wait: brands with no corpus footprint yet. If almost nobody writes about you, there is little for a model to synthesise, and customers and coverage have to come first.
What does a generative engine optimization agency do?
A generative engine optimization agency works to get a brand retrieved, cited and accurately described inside AI-generated answers — ChatGPT, Perplexity, Google AI Overviews, Gemini, Claude and the other answer engines — by making that brand findable to AI crawlers, parseable as clean extractable passages, and corroborated as a coherent entity across the third-party sources those systems draw on. The same discipline trades under four names: GEO (generative engine optimization), AEO (answer engine optimization), LLM SEO, and AI visibility. Buyers arrive using all four, suppliers pick whichever one their competitors are not using, and there is no meaningful difference in the work behind them. No answer engine exposes a ranking API or sells placement in an organic answer, so nothing in this discipline can be guaranteed by anyone — a supplier promising a citation is promising something the platform does not offer them to sell.
Our engagements here are scoped and quoted after a call, not sold from a price list. The work is sized by how many buying questions you need to cover, how badly your entity is described on sources you do not own, and how much of the fix is content versus earned coverage — and those three inputs vary by an order of magnitude between a single-product SaaS and a fourteen-market brand with a stale Wikipedia entry.
The four labels exist because three different groups named the same problem at roughly the same time and none of them yielded. Search people called it LLM SEO because it looked like SEO. Content strategists called it answer engine optimization because the output is an answer rather than a list. Academics and then vendors called it generative engine optimization because the answer is generated rather than retrieved intact. Brand and comms teams called it AI visibility because what they care about is whether the brand shows up at all. Used loosely they are synonyms, and this page treats them as one discipline throughout.
| Label | Who tends to use it | What it emphasises | Is the work different? |
|---|---|---|---|
| GEO — generative engine optimization | Agencies, analysts, the trade press | The answer is synthesized from several sources, not returned intact | No |
| AEO — answer engine optimization | Content and SEO teams | Winning the answer rather than the click | No — overlapping origin in featured-snippet work |
| LLM SEO | In-house search teams, technical marketers | Continuity with the existing search programme | No — the framing carries the most baggage |
| AI visibility | Brand, comms and executive stakeholders | Presence and accuracy of description, measured over time | No — it names the measurement, not the method |
What does the work actually consist of?
Generative engine optimization consists of three jobs that have to be done together: making your pages retrievable by the crawlers that feed answer engines, making the content on them extractable as self-contained passages, and making your brand corroborated as an entity on sources you do not control. Skipping any one of the three produces a programme that looks busy and moves nothing, and the third is the one most in-house teams under-fund because it does not live in the CMS.
- Findable. You cannot be cited by an engine you have blocked. That means an explicit robots.txt decision on the OpenAI crawlers —
GPTBotfor model training and refresh,OAI-SearchBotfor the real-time search index behind browsed answers, andChatGPT-Userfor user-initiated live fetches — plus the equivalent decision for every other engine you want to appear in. It also means server-side rendering or progressive enhancement on pages that matter, because extractors read HTML rather than a JavaScript-built DOM. Ifcurlagainst your URL returns an empty shell, you are invisible to the retrieval layer regardless of how the page looks in a browser. - Parseable. A generated answer is assembled from passages, so the unit of optimization is the passage, not the page. Question-shaped headings, a direct answer in the first forty to sixty words under each one, tables for anything comparative, explicit dates, and sentences that survive being lifted out of context. A paragraph that opens “this means that” cannot be quoted standalone and will lose to one that can.
- Citable. The engine forms its impression of your brand mostly off your own site. Published analysis of brand citations has found that roughly 85% of brand mentions come from domains the brand does not own, and that brands are around 6.5× more likely to be surfaced through third parties than through their own content. That makes review platforms, comparison content you did not write, community discussion, trade coverage and consistent entity facts across external profiles the majority of the job rather than a nice-to-have.
There is a fourth job that nobody sells and everybody needs: entity consistency. If your company name, founding date, category description and product list disagree across your website, your review profiles, your funding databases and your LinkedIn page, a model has less confidence about who you are, and options it cannot resolve cleanly drop out of shortlists first. This is dull, cheap work with a long tail of effect, and it is usually the first thing we find. Our guide to how brands appear in ChatGPT responses covers the selection mechanism underneath all of this in more depth than a service page can.
Is a GEO agency just doing SEO and digital PR under a new name?
Partly, and any supplier who denies it is selling. Technical crawlability, information architecture, structured data, content quality and earned third-party coverage are all pre-existing disciplines, and a competent GEO programme uses all of them. What is new is the target and the arithmetic: you are optimizing a brand entity for inclusion in a synthesized answer that names three to five options, rather than optimizing a URL for a position on a page that lists ten. That changes which work matters most, which is why a GEO programme that is genuinely GEO reallocates budget away from page-count content and toward corroboration, entity hygiene and passage-level rewriting of the pages you already have.
The distinction determines what you should expect to buy, what you should refuse to pay for, and how you should judge a supplier after six months. The next section takes it apart properly.
Who should not hire a GEO agency yet?
Three situations make this engagement premature. If nothing about your category is ever asked of an assistant — hyper-local trade services, pure impulse retail, categories where the buyer never researches — the prompt set you would track is thin and the programme has nowhere to compound. If your site is not indexable or not indexed at all, the retrieval layer has nothing to retrieve and the first fix is technical rather than generative. And if you need attributable revenue inside one quarter, this is the wrong channel to buy: entity signals move on a timescale of months in both directions, and the traffic that arrives directly is small enough that section three states the number plainly rather than hiding it.
We would rather tell you to wait than sell you a programme you cannot judge. The scoping call ends in a written scope and a quote, or in a recommendation to do something else first.
GEO, AEO and SEO: what is actually different
The difference between SEO and generative engine optimization is that SEO optimizes a page to be ranked among a list of results a person then chooses from, while generative engine optimization optimizes a brand and its supporting corpus to be retrieved, trusted and named inside an answer the engine writes on the user’s behalf. Ranking is ordinal and observable; citation is binary, assembled from several sources at once, and observable only by sampling. Roughly half of the underlying craft carries over unchanged. The half that does not is the half that determines whether a supplier is doing this work or relabelling the work they already sold you.
What genuinely carries over from SEO?
Five things carry over from SEO to GEO without modification, and a GEO programme that ignores them fails on the fundamentals before it reaches anything generative.
- Crawlability and indexation. Browsed AI answers retrieve against a search index. If your pages are not in an index the engine can reach, no amount of passage rewriting matters. The robots.txt decision now has more rows in it, and blocking a crawler you meant to allow is a self-inflicted wound that takes weeks to notice.
- Site health. Server-rendered HTML, working canonicals, sane internal linking, pages that return 200 and load without a consent wall in front of the content. Extractors are less patient than browsers, not more.
- Structured data. Schema.org markup does not make an engine cite you, but it makes your facts machine-readable and consistent, which is exactly the problem the entity layer is trying to solve.
FAQPage,Organization,Productand an honestdateModifieddo more work here than they ever did for blue links. - Authority and corroboration. Whatever you called link building, its underlying purpose — being referenced by sources other people already trust — is more central in GEO than it was in SEO, not less. The mechanism changed from link equity to entity corroboration; the budget line is the same one.
- Content quality. Specific, first-hand, dated, verifiable content still wins. The bar rose, because a synthesizer reading five sources will preferentially use the one that says something the other four do not.
What is genuinely new?
Four things are genuinely new in generative engine optimization, and each of them breaks a habit that SEO trained into marketing teams over twenty years.
- Retrieval and synthesis replace ranking. The engine decomposes one question into several sub-queries, retrieves against each, discards most of what it pulls, and writes a single answer from what survives. Published analysis of 548,534 pages across 15,000 prompts found ChatGPT cites roughly 15% of the pages it retrieves — the other 85% are read, evaluated and dropped without ever appearing. Being retrieved is necessary and nowhere near sufficient.
- You are cited, not clicked. The commercial event is your name appearing in the answer with your description attached, and a click is an optional follow-on that most users never make. That inverts the whole measurement model, and section three deals with the consequence honestly.
- Position has no meaning. An answer that names four brands drawn from eleven retrieved sources has no position two. Order of mention is a weak directional signal at best — a brand named first across many runs generally holds a stronger entity position than one named last — but it is not a rank, it is not stable, and any supplier reporting “your AI ranking improved to #3” is reporting a number the system does not produce.
- Framing matters as much as inclusion. Being described wrongly is worse than being absent. A model that names you as the budget option, or attributes a feature you removed two years ago, or places you in an adjacent category, is doing active commercial damage at scale — and it is doing it from stale third-party sources rather than from your website. Sentiment and description accuracy are therefore first-class objectives in a GEO programme, not a reporting garnish. Our guide on why ChatGPT recommends your competitor covers the displacement and misdescription cases in detail.
How do the three actually compare?
Used precisely, SEO, AEO and GEO name three different optimization targets: a ranked page, an extracted answer, and a synthesized answer assembled from many sources. Used commercially, AEO and GEO are the same offer with different marketing. The table below is the precise reading — it is the one worth holding a supplier to, because it names the unit of success each is actually working toward.
| SEO | AEO — answer engine optimization | GEO — generative engine optimization | |
|---|---|---|---|
| What it optimizes | A URL, for ranking against a keyword in an index | A passage, for being lifted verbatim as the direct answer | A brand entity plus the corpus that describes it, for inclusion in a generated answer |
| Unit of success | Position, then click | The answer slot — one source wins it | A citation or named mention, alongside three to five others, described accurately |
| How it is measured | Rank tracking against an index that can be queried directly, plus Search Console clicks and impressions | Whether the snippet is yours, plus impressions on the answer-bearing query | Sampling only. A fixed prompt set run repeatedly; citation rate, share of voice, source URLs cited, sentiment |
| What a win looks like | Position 8 to position 3, and sessions rise accordingly | Your paragraph is the one shown above the results | Your brand is named in most runs of the prompts you care about, described correctly, with your URL among the cited sources |
| Feedback speed | Days to weeks, against a stable number | Days to weeks | Weeks to months, against a noisy number that must be trended rather than read |
Where does a brand get lost between the question and the answer?
A brand can be dropped at five distinct points between a user typing a question and an assistant printing an answer, and the fix is different at every one. Diagnosing which stage you are losing at is the first hour of any competent engagement, because the four other fixes are wasted effort if you are failing at stage one.
| Stage | What has to be true | How a brand is lost here | What actually fixes it |
|---|---|---|---|
| 1. Entity | The engine can resolve your brand name into one coherent, known thing in a category | Name collisions, inconsistent descriptions across profiles, no presence on the sources the model treats as reference material | Entity hygiene across external profiles, consistent one-line description everywhere, structured data on your own properties |
| 2. Validation | Claims about you are corroborated somewhere other than your own marketing | Every claim traces back to your own site; no reviews, no independent coverage, no third-party comparison | Review presence, earned coverage, expert commentary, community discussion — the 85%-off-domain problem |
| 3. Retrieval | Your pages are crawlable, indexed, and rank for the sub-queries the fan-out actually issues | Blocked crawler, JavaScript-only rendering, or no content aimed at the comparison-shaped sub-queries a buying question decomposes into | Crawler allowances, server-side rendering, and content built for “best X for Y”, “X alternatives”, “X vs Y” sub-queries |
| 4. Synthesis | Your passage is usable, self-contained and adds something the other retrieved sources do not | Retrieved and discarded — roughly 85% of retrieved pages are, per the published analysis above — because the passage is vague, hedged, undated or duplicative | Passage-level rewriting: direct answers, specifics, dates, tables, original data |
| 5. Output | You are named, described accurately, and ideally linked | Used but not credited; or named with a stale, wrong or unflattering description | Correcting the upstream sources the description is drawn from; publishing the canonical facts in extractable form |
Can a brand be used without being cited?
Yes — an engine can read your page, use what it says, and never name you, and that gap is where most of the confusion in this category lives. Synthesis does not require attribution. A model that learns from your comparison table can reproduce its substance in an answer that cites three other domains, or none. The published finding that only around 15% of retrieved pages are cited is the same phenomenon counted from the other end: the other 85% were used to form the answer and then dropped from the visible source list.
Three consequences follow, and each one changes what you should be tracking. First, influence and citation are different metrics — your content can be shaping answers while your citation rate reads zero. Second, a brand mention without a link is still commercially valuable, because the user asked for a recommendation and got your name; measuring only referral traffic misses this entirely. Third, being used without being named is not a bug to appeal, it is how synthesis works, and the response is to strengthen the entity signals that make naming you the natural choice rather than to try to force attribution.
This is also the honest answer to “why did my competitor get named when my page is clearly better?” The answer being written is not a judgement about your page. It is an accumulated impression of your brand as an entity, grounded in retrieved passages after the fact.
The diagnostic is cheap and worth running yourself before you brief anyone. Take one buying question, ask it of an engine that shows its sources, and check two separate things: whether your domain appears in the source list, and whether your facts appear in the prose. Four outcomes fall out of that, and each has a different fix. Named and cited means the system is working and the job is holding it. Cited but not named means your page was useful and your entity was not strong enough to be worth mentioning — a corroboration problem. Named but not cited means your entity is known and your own pages are not the evidence, so the description you are getting is coming from sources you do not control and should go and read. Neither means you are being lost upstream, at retrieval or before it, and passage rewriting will not save you. Run it across ten questions and the pattern is usually unambiguous within an hour.
The one red flag that settles it
If a supplier’s GEO deliverable list is identical to their SEO deliverable list, it is SEO with a new label. Put the two proposals side by side. If the GEO scope is a technical audit, keyword research, four blog posts a month and a link-building allocation — the same scope, the same hours, at a higher rate because the acronym changed — you are buying a rebrand. A real GEO scope contains line items that have no SEO equivalent: a defined prompt set built from your buyers’ actual questions, competitor share-of-voice benchmarking on that set, entity reconciliation across sources you do not own, description-accuracy remediation on stale third-party profiles, and passage-level rewriting judged by extractability rather than by word count. Ask for the prompt set before you sign. A supplier who cannot show you one has no measurement plan, and a supplier with no measurement plan cannot tell you whether anything they did worked.
The reverse failure is worth naming too, because it is the more expensive one. A supplier who treats GEO as entirely new, discards the search programme, and starts publishing “AI-optimized” content into a site that is not properly indexed has removed the foundation the retrieval stage depends on. The correct position is that GEO is a reallocation of an existing budget with new priorities and a new measurement layer, not a greenfield channel. Our guide on how to appear in ChatGPT answers sets out the four signals that reallocation should be aimed at.
What can and cannot actually be measured
No answer engine exposes a ranking API, an impression count, or any endpoint that reports where a brand stands in its answers — so every AI-visibility number in this category, including ours, is produced by asking the engine questions and recording what it says. There is nothing to query for position, which means there is nothing to guarantee and nothing to audit against a platform-issued source of truth. This section comes before any description of what we sell, deliberately, because a buyer who does not understand the measurement cannot evaluate the offer, and the measurement is where this category does most of its lying.
Why can nobody guarantee a citation?
Nobody can guarantee a citation in an AI answer because no engine sells organic placement, exposes a ranking parameter, or produces a deterministic output. The same prompt, issued twice, can return different brands. Answers vary by user, by session, by conversation history, by model version and by whether the engine chose to browse at all for that particular question. A guarantee would require control over a system whose operator does not offer that control to anyone at any price.
This is different from the situation in paid channels, and the difference is worth stating precisely because buyers move between the two. In advertising you buy an auction entry with published mechanics; in generative engine optimization you improve the probability of an outcome you cannot observe directly. Any supplier offering guaranteed citation, guaranteed AI ranking or a “#1 in ChatGPT” position is describing something the platform does not sell. We do not offer it, and we would treat anyone who does as disqualifying.
How do AI visibility tools actually get their numbers?
Every AI-visibility tool samples rather than observes: it runs a set of prompts you or it defined, records which brands and URLs appear, and repeats. None of them see what real users asked, what real users were shown, or how often. The category splits into two technical approaches and both are sampling — the difference is what they sample.
| Approach | How it works | What it actually reflects | Where it misleads |
|---|---|---|---|
| API sampling | Prompts are sent to the model provider’s API and the responses are parsed for brand and URL mentions | API-layer responses, at whatever model version and configuration the tool selected | The API response is not the consumer product. Browsing behaviour, system prompts, personalization and product-surface differences mean the answer a customer sees in the app can differ from the answer the API returns |
| UI scraping | An automated session drives the consumer interface and captures the rendered answer | Something much closer to a real user’s answer, from one synthetic session | Sessions are logged-out or tool-owned, carry no real memory or history, and are fragile to interface change; blocking and rate limits shape the sample |
| Both | A defined prompt list, run on a schedule, results aggregated into rates | A synthetic panel you chose, tracked consistently | Presenting a sampled rate as though it were an observed share of real user answers |
Neither approach observes end-user behaviour, and no vendor in this category has privileged access that changes that. Two tools running different prompt sets against different model versions will disagree, and they are not supposed to agree — the number is a property of the methodology as much as of your brand. The practical implication is that you should never compare a citation rate from one tool against a citation rate from another and conclude anything. Compare a tool against itself, over time, on a prompt set that has not changed.
How volatile are the numbers, honestly?
Volatility in AI-visibility measurement is severe enough that any single reading is close to meaningless. Published synthesis in this category reports that only around 30% of brands remain visible from one run of the same prompt to the next, and our own guide on how to track ChatGPT visibility records comparable instability across major model updates. Answers vary by user, by session and by model version, and a model release can reshuffle a category’s shortlist in a week for reasons no supplier can trace.
Three operating rules follow from that, and they are the difference between a dashboard that informs decisions and one that generates panic.
- Never report a single run. A citation rate means the proportion of runs across a prompt set in which you appeared, aggregated over a period. One prompt, one run, one screenshot is an anecdote — useful for a stakeholder conversation, worthless as a metric.
- Never change the prompt set silently. The panel is the instrument. Adding fifteen prompts you are likely to win on will lift your citation rate without changing anything about your brand. When the set changes, version it and report both panels in parallel for a period.
- Never read it daily. Run-to-run variance is larger than any weekly movement you are likely to cause, while the underlying signals — third-party mentions, corroboration, entity consistency — move on a timescale of months. Weekly capture, monthly interpretation.
How much traffic does this actually send?
Referral traffic arriving directly from LLM interfaces is currently tiny: third-party analyses put it at roughly 0.1% of website traffic. That is the strongest rational objection to this entire category, it is usually buried on the fourth page of a competitor’s deck if it appears at all, and it belongs here, before the sales material, because a buyer who has not priced it in will be unhappy in month four.
Two things are true at once, and holding both is the whole judgement call. Direct referral volume is small and will stay small for as long as answers satisfy the question without a click. And the same interaction can decide a shortlist before any measurable session exists — the buyer asks for options, gets three names, searches one of them by brand, and arrives through a channel your analytics labels organic or direct. That is why branded-search lift and self-reported attribution belong in the reporting stack, and it is also why nobody in this category can hand you a clean ROI number. If your business case requires attributable last-click revenue from AI referrals inside two quarters, the honest answer is that this channel will not produce it, and you should spend the money somewhere it can.
So what can be measured honestly?
Five things can be measured honestly, and the useful ones are all rates on a fixed prompt set rather than absolutes. Each has a defensible definition and a stated limit, and a supplier who will not state the limit alongside the number should not be trusted with the number.
- Citation frequency on a defined prompt set. The proportion of runs, across a versioned panel of buying questions, in which your brand is named. Honest because the denominator is explicit. Limited because you chose the denominator.
- Share of voice against named competitors. Your citations divided by the total citations of a specific competitor set on the same prompts in the same runs. This is the most decision-useful metric in the category, because it is relative — a model update that drops everyone shows up as flat share of voice rather than as a fake crisis.
- Which source URLs get cited. The engine usually shows its sources, and this is the closest thing to a diagnostic the category has: it tells you whether you are being cited through your own pages, through a review platform, through a comparison article you did not write, or not at all. It is also how you find which off-domain sources describe you wrongly.
- Sentiment and description accuracy. Whether the answer describes your product, pricing model and category correctly. Scored by reading the answers, because there is no automated substitute for judgement here. This is frequently the highest-value finding in a first month of tracking and the one no dashboard surfaces on its own.
- Directional proxies. Branded-search volume, direct-traffic movement, assisted conversions, and self-reported attribution on your forms — the “how did you hear about us” field, which is the only measurement instrument in this section that hears from actual buyers. None of these prove causation. Together, moving in the same direction as your citation rate, they build a case.
These are sampled estimates on a prompt set you chose, tracked consistently over time. That makes them useful as a trend and misleading as an absolute.
That sentence is the honest framing, and it should appear on the front page of any report you are given. A citation rate of 34% does not mean 34% of people asking about your category are shown your brand — it means that on your panel, on the runs that were executed, at the model versions that were used, you appeared about a third of the time. Read as a level, it is fiction. Read as a series it can be real evidence — illustratively, a panel that moves from 18% to 34% over five months without the prompt set changing, while share of voice against the same four named competitors also rises, is a defensible claim that something worked. The full metric definitions are in our guide to ChatGPT visibility metrics and how to build a defensible weekly dashboard.
Where does measure.co fit, and what are its limits?
measure.co is our own AI-visibility tracking product, from $89/month, covering ChatGPT, Gemini, Claude and Perplexity: citation rate, share of voice, competitor benchmarking, a prompt-coverage dashboard, weekly reports and alerts. It is the measurement layer under a GEO engagement, and clients can and do run it without buying anything else from us.
It has exactly the same structural limits as everything else described in this section. It samples rather than observes. It runs the prompt set you and we defined, not the questions your actual buyers asked. It is subject to the same run-to-run volatility, and it will disagree with any other tool you point at the same brand. We do not have access to data nobody else has, and we will not claim otherwise to sell a subscription. What it gives you is a consistent methodology applied on a consistent cadence to a panel you can inspect and version — which is the only thing that makes a trend in this category trustworthy, and it is a genuinely useful thing to own.
Which claims in this category should you discount?
Five claims appear in nearly every GEO pitch deck in this category, and each one is either a sampled estimate presented as an observation or an outright impossibility. The table below is what sits behind each, and how much weight to give it.
| Claim commonly made | What is actually behind it | How much weight to give it |
|---|---|---|
| “We guarantee you will be cited in ChatGPT” | Nothing. No engine sells organic placement or exposes a ranking control | None. Treat it as disqualifying |
| “Your AI ranking improved to position 3” | Order of mention within one answer, in a sample of runs. There is no ranking to hold a position in | Very little. Weak directional signal at best; never a rank |
| “You appear in 62% of AI answers in your category” | You appeared in 62% of runs on a prompt panel somebody chose, at the model versions sampled | Some — only if the panel is disclosed, versioned and unchanged between reports |
| “We have direct data from the AI platforms” | API sampling or UI scraping. There is no visibility data feed to have | None. Ask which of the two it is; the answer tells you whether they understand their own product |
| “AI search will replace Google, so this is urgent” | Genuine growth in assistant usage, alongside LLM referral traffic at roughly 0.1% of website traffic per third-party analysis | Some — as a reason to start early and cheaply, not as a reason to reallocate a search budget wholesale |
| “Our citation rate is higher than the other tool’s” | Two different prompt sets, sampling methods and model versions producing two different numbers | None. Cross-tool comparison is not meaningful; compare a tool against itself over time |
What to demand from any supplier before you sign
Ask for five things, in writing, before money changes hands. One: the prompt set. The actual list of questions they will track, built from your buyers’ language, with a stated version and a rule for how it changes. Two: the sampling method. API or interface, which engines, which model versions where knowable, how many runs per prompt, on what cadence. Three: a named competitor set for share of voice, agreed before the baseline is taken rather than chosen afterwards to flatter the chart. Four: the baseline itself — measured and reported before any work starts, including the answers where you are described inaccurately. Five: a written statement of limits, in their words, saying that the figures are sampled estimates on a chosen prompt set and that no citation can be guaranteed. A supplier who supplies all five may still not be the right one for you, but a supplier who will not supply them cannot demonstrate that anything they do works, and neither can you.
How an answer engine decides what to cite
An answer engine decides what to cite in two stages: it retrieves a small set of documents that look relevant to the question, then synthesizes an answer out of what those documents say and attaches citations to the passages it actually leaned on. Every lever you have sits inside one of those two stages, and the two fail independently — a page that is never retrieved cannot be cited however well it is written, and a page that is retrieved but written as an unbroken slab of marketing prose gives the synthesizer nothing clean to lift. What follows is what is observable from outside the system. No engine publishes its selection function, and any agency that claims to have reverse-engineered one is describing a model of something it cannot see.
The engines in scope for most commercial briefs are ChatGPT’s browsing and search modes, Perplexity, Google’s AI Overviews and AI Mode, Microsoft Copilot, Gemini and Claude. They differ in index, in interface and in how aggressively they attribute, but they share the same two-stage shape. That shared shape is why one programme can serve all of them, and why a fix at the retrieval layer helps everywhere at once while a fix at the synthesis layer helps only where the engine attributes visibly.
What happens between the question your buyer types and the answer they read?
Five things happen between the question and the answer, and your content is only eligible at three of them. First, the engine decides whether to browse at all or to answer from what the model already carries. Second, it rewrites the buyer’s question into a set of search queries — Google has publicly described this expansion as query fan-out, and the practical consequence is that the query actually run against the index is rarely the sentence the buyer typed. Third, it retrieves candidate documents from an index; for ChatGPT’s real-time retrieval that index leans on Bing, as our guide to how brands appear in ChatGPT sets out. Fourth, it selects a handful of those candidates to place in the model’s context window. Fifth, it writes an answer and attaches citations.
The fifth step is the one operators consistently miss. Citations are attached at synthesis time, not at retrieval time. Being in the retrieval set is necessary and not sufficient; the model can read your page, use what it learned, and cite somebody else — or nobody. That is why the work splits into three hurdles rather than one, and why an agency that only offers “more content” is optimizing for a single hurdle out of three.
| Hurdle | What it asks | How it typically fails | How you test it |
|---|---|---|---|
| Retrievable | Can the engine fetch and index this URL at all? | A crawler blocked in robots.txt; a bot manager returning 403 to every agent that is not Googlebot; content that only exists after client-side JavaScript; the page not present in the underlying index. | Fetch the URL with each crawler’s user-agent string and read the status code and the raw body. Check server logs for the agents you expect. |
| Parseable | Can a clean, self-contained span be extracted from it? | Answers buried in the fourth paragraph; headings that label a topic rather than pose a question; sentences that open with “This means” or “It does this by”; a definition split across three paragraphs. | Take any 60-word window from the page and read it cold. If it needs the sentence above it to make sense, it is not a passage. |
| Quotable | Is there a specific, attributable claim worth pointing at? | Copy that asserts quality rather than fact; claims that a hundred other pages also make; superlatives with no source; nothing a citation could usefully anchor to. | Ask what fact a reader would lose if this passage vanished. If the answer is “nothing”, there is nothing to cite. |
Why does passage structure matter more than page structure?
A model lifts a span, not a document, so the unit that has to be well made is the passage rather than the page. Retrieval systems chunk documents before they embed or rank them, and the chunk boundary rarely respects your argument. A section that builds carefully across four paragraphs to a conclusion is a section whose conclusion arrives in a chunk that no longer contains the reasoning — and a chunk that cannot stand alone is a chunk that will not be quoted.
Three properties make a passage survive extraction. The subject is explicit in every sentence, so the noun the reader needs is present rather than implied. There is no anaphora across the opening boundary — no “this”, “that”, “it” or “the above” pointing backwards out of the chunk. And the definitional sentence sits first, before the qualification, because an extractor reading top-down finds the answer before it finds the caveats.
Headings do double duty here. They are a human navigation aid and they are also a retrieval anchor: a heading phrased as the question a buyer would type gives the engine an explicit match between the query it fanned out and the span sitting under that heading. A heading that reads “Our approach” anchors nothing. A heading that reads “How long does a freight invoice audit take?” anchors a specific query and tells the extractor exactly where the answer begins. This is not a formatting preference. It is the difference between a page that has an answer in it somewhere and a page from which an answer can be lifted.
Why does a brand get used without being cited?
A brand gets used without being cited when the engine takes a fact from the page but has no reason to point at the page as the fact’s source. Citation is a pointer to provenance, so the question the synthesizer is implicitly answering is not “which page was helpful” but “which page is where this particular claim comes from”. Content that restates what is already common across the corpus is helpful and unattributable at the same time, because there is no single origin worth naming.
Four patterns account for most of the gap between being read and being named. The claim is generic, so any of a dozen sources could have supplied it. Another page states the same thing more crisply, so it wins the pointer on extractability alone. Your page is a secondary restatement of a fact that visibly originates elsewhere — you cite a study, so the engine cites the study. Or your page is one of several near-identical pages on your own domain, and the engine picks one and drops the rest. The remedy in every case is the same: be the place a specific, checkable claim originates. A methodology you name and describe, a figure you measured and dated, a definition you wrote, a comparison table you built — these are citable because they have a source, and the source is you. We take the same view of why a competitor gets named instead of you, which our guide to why ChatGPT recommends competitors works through in detail.
Why does the rest of the web outweigh what your own site says about you?
Models reflect what a corpus broadly says, so a claim that appears only on your own domain carries less weight than a claim repeated across independent sources. This is the single hardest thing to explain to a marketing team, because it inverts the usual logic of a website. You control your own copy completely, which is exactly why it is the weakest evidence available: a self-description is a claim about the claimant, and the engine has read ten thousand sites that all describe themselves as the leading provider in their category.
Corroboration is what converts a self-description into something an engine will repeat. If your own site says you serve mid-market distributors, and a trade publication, two industry directories, a conference speaker bio and a partner’s integration page all say the same thing in their own words, the claim has become a property of the corpus rather than a property of your homepage. If your site says you are the market leader and nothing else does, the claim is invisible — not disbelieved, just unsupported and therefore unusable.
Two practical consequences follow. First, superlatives are wasted words in this channel; replace them with facts that can be independently checked, because a checkable fact can be corroborated and an adjective cannot. Second, the order of work is fixed by this dependency: the corroboration layer determines what the content layer is allowed to claim, which is why the workstreams below are sequenced the way they are rather than started in parallel.
How much does recency actually matter?
Recency matters in proportion to how fast the subject moves, and it matters more in browsed answers than in answers the model gives from what it already carries. For a stable definitional question, a well-made page from two years ago competes fine. For anything with a price, a version number, a regulation, a platform feature or a date attached, staleness is disqualifying — an engine summarizing a moving subject will prefer a source that looks current, and a page describing last year’s state of the world is worse than no page because it is confidently wrong.
The observable freshness signals are ordinary: a visible last-updated date in the page body, dateModified in your structured data, lastmod in the XML sitemap, and content that has actually changed. The last item is the one that matters. Bumping a timestamp on a page whose text has not moved is freshness theatre, it is trivially detectable by anything that diffs content between crawls, and we do not do it. What we do instead is maintain a review cadence per page based on how volatile its subject is, and record what changed each time, so that the date on the page is a statement about the content rather than a decoration.
What is known here, and what is inference?
Almost everything about answer-engine selection is inference, and the honest division runs roughly as follows. Published and first-party: crawler user-agent tokens and how robots.txt directives are interpreted, the snippet controls that suppress text from generated answers, structured-data documentation, and Google’s statement that its AI surfaces are served by the same crawler as Search. Observable by anyone willing to do the work: which URLs get cited for which prompts, repeatedly, across time and across engines; whether a change to a page precedes a change in citation; whether an engine can fetch a URL at all. Inference: everything about weighting, thresholds, and why one source beat another for a particular answer.
The distinction is not academic, because it determines what an agency can honestly sell. We can promise a method, a cadence, a measurement discipline and a written record of what changed and what happened next. We cannot promise a citation, a position, or a share of answers, and nobody can — the selection happens inside a system none of us can inspect, and it is re-run from scratch on every prompt.
What we will not claim about this channel
We will not claim a deterministic model of any engine’s ranking. We will not promise that your brand will be cited, or cited by a given date, or cited for a given prompt. We will not present prompt-level citation counts as a stable metric when the same prompt can return different sources on consecutive runs. What we will do is baseline a fixed prompt set, re-run it on a fixed schedule, record the results with dates, and be explicit about which movements are large enough to mean anything and which are noise.
The five workstreams
A generative engine optimization programme is five workstreams run in a fixed order: crawl and access, entity foundation, answer-ready content architecture, citation and authority engineering, and continuous experimentation and monitoring. They are layers rather than a checklist, because each one gates the value of the ones above it — content architecture on a site the engines cannot fetch produces nothing, and citation building for a brand the engines cannot identify produces credit for somebody else. The sections after this one open up the first three in full; the table at the end of this section is the map.
The layers overlap in calendar time and do not overlap in dependency order, which is a distinction worth being precise about. In a real engagement, crawl and access is diagnosed in the first week and usually remediated inside the first month, because it is infrastructure work with a small number of owners. Entity foundation starts in the first month and never really finishes, because half of it is other people’s surfaces on other people’s schedules. Content architecture starts once the entity source-of-truth document exists, so that every page written from that point forward asserts the same facts. Citation and authority work starts as soon as there is something worth pointing at, which in practice means after the first tranche of rewritten pages. Monitoring starts before any of it, because a baseline taken after the changes is not a baseline.
Two of these layers produce visible artifacts and three do not, and that asymmetry is where most GEO retainers go wrong. Rewritten pages and earned coverage are legible in a monthly report; a corrected bot rule, a standardized description string and a dated prompt transcript are not. An agency paid to produce a report will drift toward the legible layers, which is why we write the sequencing into the engagement rather than the deck, and report the unglamorous work explicitly — what was fetched, what returned what status code, which strings changed on which profiles, and what the engines said before and after.
Layer 1: Crawl and access
Crawl and access asks one question: can the engines fetch you at all, and does what they fetch contain your content. It covers robots.txt decisions for each named crawler, the bot-management and WAF rules that block agents without ever touching robots.txt, server-side rendering versus client-side JavaScript, HTTP status codes, redirect chains, canonical consistency, sitemap accuracy, and the snippet directives that let a fully crawlable page suppress its own quotability. It is trying to fix the failure where a brand invests in content for a year while a single line in a config file, or a security product doing exactly what it was bought to do, has removed the site from consideration entirely. This layer is cheap, fast, and the only one where a single defect can zero out everything else.
Layer 2: Entity foundation
Entity foundation asks whether the engine knows what your brand is, unambiguously, and can tell you apart from everything with a similar name. It covers the exact strings you use for your name, description, category and core facts across every surface you control; Organization, Person, Product and Service structured data with accurate sameAs links; third-party profiles and industry directories; and active disambiguation when an engine has merged you with someone else or is still describing you under a name you retired. It is trying to fix the failure where the citation goes to a competitor with a similar name, or where the engine describes your company using facts from a different company, or where a rebrand happened eighteen months ago and the corpus has not noticed. Entity work is slow and unglamorous, and it is the layer that decides whether authority built anywhere else accrues to you.
Layer 3: Answer-ready content architecture
Answer-ready content architecture asks whether your content is shaped so a span can be lifted out of it. It covers page-level information architecture organized around the questions buyers actually ask rather than keyword variants, passage-level rewriting so that every section answers its own heading in its first sentence, the formats that get quoted most readily — definitions, comparison tables, stepwise procedures, FAQ blocks — specificity discipline, and a maintenance cadence set by how fast each subject moves. It is trying to fix the failure where the site ranks perfectly well, gets retrieved, and still never appears in an answer because there is no clean, self-contained, specific span anywhere in it. Most of the writing budget goes here, and it is also where the most money is wasted when layers 1 and 2 have been skipped.
Layer 4: Citation and authority engineering
Citation and authority engineering asks whether independent sources corroborate what your own site says. It covers earned coverage in publications your buyers and the engines already read, expert contribution and commentary where your named people say something worth quoting, original research and data you publish so others have a reason to cite you, review platforms and industry directories that are genuinely relevant to your category, partner and integration pages, conference and podcast appearances that produce indexable transcripts, and consistency of your core facts across all of it. It is trying to fix the corpus-consensus problem: a claim that lives only on your domain is a claim the engine has no reason to repeat.
This is also the layer where the category’s bad practices live, so we are explicit about what we decline. We do not create or commission fake reviews. We do not post to Reddit, forums or communities under manufactured identities, or pay others to. We do not buy placements and present them as independent coverage. We do not create or edit Wikipedia or Wikidata entries for clients. Beyond the ethics, these tactics fail on their own terms: the mechanism you are trying to exploit is corpus-wide agreement, and a cluster of near-identical mentions appearing in a short window is precisely the pattern that reads as manufactured to both platform moderation and to anyone checking the sources by hand.
Layer 5: Continuous experimentation and monitoring
Continuous experimentation and monitoring asks whether you know if any of it worked. It covers a fixed prompt set representing how your buyers actually ask, baselined before changes and re-run on a schedule; recording answers verbatim with dates rather than only a score, because the wording of an answer tells you more than its presence or absence; referral traffic from answer-engine domains in your analytics; server-log analysis of which AI crawlers fetch which URLs and how often; and a change log that lets you line up what you shipped against what moved. Our own tracking product, measure.co, exists for the recurring half of this — citation rate, share of voice, competitor benchmarking and prompt coverage across ChatGPT, Gemini, Claude and Perplexity.
Two honesty constraints govern this layer. Answer-engine outputs are not deterministic, so the same prompt can return different sources on consecutive runs and a single-run comparison proves nothing. And attribution is worse here than in paid channels: a buyer who reads about you in an answer and searches your brand name the next day arrives as direct or organic traffic with no trace of the answer that sent them. We plan for that by pairing engine-side measurement with a self-reported attribution question at the point of enquiry, which is the only place the buyer can tell you what actually happened. Our guide to how buyers evaluate vendors inside ChatGPT covers what that journey looks like from the buyer’s side.
| Layer | Typical failure it addresses | What changes, on the site or off it | How you would know it worked |
|---|---|---|---|
| 1. Crawl and access | The engines cannot fetch the site, or fetch an empty shell, or are blocked by a bot manager nobody in marketing knows about | On site: robots.txt, WAF and bot-management rules, rendering strategy, status codes, redirects, canonicals, sitemap, snippet directives | Named crawler user-agents appear in server logs with 200 responses; a raw fetch of key URLs returns the full content; AI-crawler hits grow rather than flatline |
| 2. Entity foundation | The engine cannot say what the brand is, confuses it with a similar name, or repeats facts from a retired identity | Both: name, description, category and core facts standardized everywhere; Organization, Person, Product, Service and sameAs markup; third-party profiles corrected | A “what is brand?” prompt run across engines returns a description that matches your own, with no merged or wrong facts; the same answer is stable over repeated runs |
| 3. Content architecture | Pages are retrieved but nothing quotable can be lifted from them | On site: question-shaped headings, first-sentence answers, definitions, comparison tables, procedures, FAQ blocks, specificity, maintenance cadence | Cited passages start appearing verbatim or near-verbatim in answers; the URLs cited are the pages you rewrote; referral traffic from engine domains appears in analytics |
| 4. Citation and authority | Every claim about the brand originates on the brand’s own domain, so the corpus has no independent agreement to reflect | Off site: earned coverage, expert contribution, original research, relevant directories and review platforms, partner pages, indexable transcripts | Independent URLs stating your core facts appear in engine answers alongside or instead of your own; brand mentions grow on domains you do not control |
| 5. Experimentation and monitoring | Nobody can say whether the previous four layers changed anything | Process: baselined prompt set, scheduled re-runs, verbatim answer logging, log analysis, referral tracking, change log, self-reported attribution at enquiry | You can point at a dated change and a dated movement and argue about causation with real numbers instead of impressions |
Sequencing is the whole argument
Layers 1 and 2 gate everything above them. If the engines cannot fetch your pages, no amount of rewriting is visible to them; if the engines cannot identify your brand, authority you build attaches to the wrong entity or to nothing. Skipping straight to content architecture is the most common wasted spend in this category, because it is the layer that feels most like work and produces the most output — forty rewritten pages is a visible deliverable, and a corrected bot-management rule is one line. We run crawl and access first because it takes days and can invalidate everything else, entity foundation second because it takes months and everything else accrues to it, and content third because that is the point at which the work can actually be seen.
Crawl and access: can the engines fetch you at all
Crawl and access is the layer that determines whether your content is eligible for retrieval, and blocking a crawler removes you from consideration by that engine entirely rather than merely reducing your chances. There is no partial credit here and no compensating strength: a page that returns 403 to OAI-SearchBot is not a weak candidate for a ChatGPT citation, it is not a candidate. This layer is also the fastest to diagnose and usually the cheapest to fix, which is why it runs first.
Which crawlers actually matter, and what does each one do?
The crawlers that matter split into three functions, and conflating them is the commonest source of bad robots.txt decisions. Search crawlers build the index an engine retrieves from when it browses. Training crawlers gather content used to train or refresh a model, which affects what the model can say without browsing. User-triggered fetchers load a specific page live because a person or an agent asked for it. Blocking one does not block the others, and the trade-offs are different in each case.
Two names deserve to be separated explicitly, because they are routinely treated as interchangeable. OAI-SearchBot is the agent behind ChatGPT’s search and browsing retrieval — it is your search visibility. GPTBot gathers content for model training and refresh — it is a different decision with different consequences. A publisher who blocks GPTBot to keep their archive out of training and leaves OAI-SearchBot allowed has made a coherent choice: no training use, full search visibility. A publisher who blocks both has chosen to be absent from ChatGPT answers, which is a legitimate position and should be a deliberate one.
| Crawler | What it is for | Consequence of blocking it |
|---|---|---|
OAI-SearchBot | Indexes pages for ChatGPT’s search and browsing retrieval — the search-visibility agent, distinct from training | Your pages are not in the candidate set when ChatGPT browses. This is the single most consequential block for ChatGPT citation. |
GPTBot | Fetches content used for OpenAI model training and refresh | Your content is excluded from training use. Does not by itself remove you from browsed answers, but removes you from what the model can say unprompted by a fetch. |
ChatGPT-User | Fetches a page live when a user action or an agent requires the current version of it | User-initiated visits to your pages fail. A buyer who asks ChatGPT to read your pricing page gets an error instead of your page. |
OAI-AdsBot | Reported by practitioners working on the ads side as the agent that must be allowed for ads-related functionality | Reported to affect ads-related functionality. Treat this as third-party practitioner reporting rather than an OpenAI statement — we have not found first-party confirmation, and we say so rather than presenting it as documented. |
PerplexityBot | Indexes pages for Perplexity’s answer index | Your pages are not available as Perplexity citations. |
Perplexity-User | Live fetch triggered by a specific user request | User-initiated fetches of your pages fail. |
Googlebot | The crawler behind Google Search and, per Google’s own documentation, its AI surfaces including AI Overviews and AI Mode | Removal from Google Search and from the AI surfaces built on the same index. There is no separate opt-in for AI Overviews that does not run through Search. |
Google-Extended | A robots token controlling whether your content is used for Gemini app and Vertex AI grounding and training | A training and grounding decision, not a Search or AI Overviews decision. Blocking it does not remove you from AI Overviews, and operators regularly block it believing that it does. |
bingbot | Indexes for Bing, which underpins several third-party answer products and, per our own guide, ChatGPT’s real-time retrieval | Broad and often underestimated. A Bing block propagates into answer surfaces that never mention Bing anywhere in their interface. |
ClaudeBot | Anthropic’s crawler | Your content is unavailable to Claude’s retrieval and training use. |
CCBot | Common Crawl, an open corpus used as an input by many model builders and research datasets | Exclusion from a corpus that many downstream systems draw on, including ones that do not run their own crawler. |
Applebot / Applebot-Extended | Apple’s crawler, with Applebot-Extended as the separate training opt-out token | Blocking Applebot affects Apple’s search surfaces; blocking only Applebot-Extended is a training decision that leaves search access intact. |
Crawler tokens and their documented behaviour change without much fanfare, and a table is a snapshot. Verify each agent against the engine’s published crawler documentation on the day you write the file, and then verify it again against your own server logs a week later, because the log is the only evidence that what you wrote had the effect you intended.
Is allowing every crawler the right answer?
No — robots.txt is a strategic decision with a real trade-off, and treating it as a checkbox is how organizations end up making a licensing decision by accident. Allowing retrieval means being quotable, and being quotable means an engine can answer your buyer’s question using your content without your buyer ever arriving on your site. For a business whose website is a sales asset, that trade is usually worth taking: an answer that names you is a recommendation with distribution you could not otherwise buy. For a business whose content is the product — subscription publishing, proprietary research, paid data, a course library — the same trade can be straightforwardly bad, and declining it is a legitimate commercial position rather than a failure of nerve.
There is a coherent middle, and most commercial sites should at least consider it: allow the search and retrieval agents, decline the training agents. Allowing OAI-SearchBot and Googlebot while disallowing GPTBot and Google-Extended keeps you fully present in browsed answers and out of training corpora. It is not a free lunch — content the model has not learned cannot be recalled without a fetch — but it is a defensible line, and it is one you should draw deliberately, in a meeting, with whoever owns the content strategy in the room.
Two practical rules make whatever you choose survive contact with reality. Write the file explicitly rather than relying on a permissive default, so the next person to open it can see the decision that was made. And record the reasoning in a comment in the file itself, because a bare Disallow line with no explanation will be either blindly copied or blindly deleted by whoever inherits it.
What is llms.txt, and does it work?
llms.txt is a proposed Markdown file at your site root that lists and describes your most useful pages, intended to give language models a curated route into a site instead of leaving them to infer structure from navigation. The proposal is sensible on its face: a short, clean, link-annotated index is easier for a model to consume than a rendered page full of navigation chrome, and the file costs an hour to write and nothing to maintain if your site structure is stable.
Its effect is not established, and we will not pretend otherwise. Adoption is uneven, no major answer engine has publicly committed to consuming it as a retrieval or ranking input, and we have not seen evidence that would let anyone honestly attribute a citation to its presence. We publish one on this site and we usually recommend clients do the same — not because we can show it moves anything, but because the cost is an hour and the downside is zero. Any agency presenting llms.txt as a primary lever is selling you a file, and you should ask them for the evidence before you pay for the sprint.
Does your content survive without JavaScript?
Content that only exists after client-side JavaScript may not survive retrieval, and the test takes thirty seconds. Fetch the URL the way a crawler would — a plain HTTP request with no browser engine behind it — and read what comes back. If the body is an empty <div id="root"> and a bundle reference, then whatever the page shows a human is not what an extractor receives. Some crawlers render and many do not, rendering budgets are finite and unevenly allocated, and a site that depends on being rendered has made its visibility conditional on somebody else’s spare capacity.
The fix is architectural rather than clever: server-side rendering, static generation or prerendering for any page that carries content you want quoted. Within a rendered page, the distinction that matters is whether content is present in the HTML or injected later. Text inside a collapsed accordion or an inactive tab is usually fine if it is in the served markup. Text fetched on scroll, loaded behind an interaction, or assembled from a client-side API call usually is not. Check the served HTML, not the browser’s inspector, because the inspector shows you the DOM after JavaScript has run and will tell you everything is fine.
What breaks quietly at the infrastructure layer?
The blocks that cost the most are the ones nobody chose. robots.txt is at least visible to anyone who looks; a bot-management product returning a challenge or a 403 to any agent outside its allowlist is invisible to marketing entirely, and it is the single most common accidental exclusion we find. Security teams allowlist Googlebot because someone complained about rankings once; nobody ever complained about OAI-SearchBot, so it is not on the list, and the site has been absent from ChatGPT for as long as the rule has existed.
- Status codes. Soft 404s that return
200with an error page teach crawlers that your error page is content. Intermittent5xxduring crawl windows suppresses recrawl rates. Arobots.txtthat itself returns5xxis treated conservatively by some crawlers, which can amount to a site-wide disallow you never wrote. - Redirects. Chains of three or more hops lose crawlers and drop query parameters. A redirect that lands on a URL different from the declared canonical creates a loop of contradictory signals. Redirecting every unmatched URL to the homepage instead of returning
404manufactures thousands of soft 404s. - Canonical consistency. The URL in
rel=canonical, the URL in the sitemap, the URL your internal links point at and the URL that actually serves the content should be one string. Trailing slashes, uppercase characters, tracking parameters andwwwvariants each split it. - The accidental
Disallow. A stagingDisallow: /promoted to production. ADisallow: /resources/written to hide a directory of PDFs that now sits above the entire content library. A wildcard for a URL parameter that also matches half the blog. Every one of these is a single line, and every one removes a section. - Snippet directives. A page can be perfectly crawlable and still muzzled.
nosnippet,max-snippet:0and the element-leveldata-nosnippetattribute all restrict the text an engine may display, and Google documents that these controls extend to its AI surfaces. A legal team that appliednosnippetto a template three years ago has been suppressing quotable text ever since. - Rate limiting. Aggressive limits that time out unfamiliar agents look identical to a slow site from the crawler’s perspective, and crawlers respond to timeouts by crawling less.
The diagnostic that resolves all of this is the same one every time: fetch your key URLs with each crawler’s user-agent string, from outside your own network, and read the status code, the headers and the raw body. Then look at server logs for those agents over a trailing month and count the hits. A crawler that is allowed, unblocked and never arriving is a different problem from a crawler that arrives and gets a 403, and you cannot tell which you have without the log.
Entity foundation: does the engine know what your brand is
An entity, in this context, is the thing an engine believes your brand to be — a cluster of facts (name, category, location, people, products, relationships) that the engine has assembled from everything it has read and treats as one identifiable subject. Entity ambiguity is fatal to a visibility programme because every other signal has to attach to something, and if the engine cannot tell which something you are, the authority you build attaches to the wrong cluster or to none. This layer is slow, it produces almost no visible output, and it determines whether the rest of the programme compounds or dissipates.
What does entity ambiguity actually look like?
Ambiguity shows up in four recognizable shapes, and each has a different remedy. Two companies with similar names in adjacent categories get merged, so the engine describes your software company using the staffing agency’s headcount and founding date. A brand that shares a word with a common noun — a company called Summit, Anchor, Vantage or Meridian — gets diluted, because most of the corpus mentioning that word is not about you and the engine has no strong reason to resolve toward a mid-sized business. A rebrand the corpus has not caught up with leaves the engine answering under a name you retired eighteen months ago, usually with the old positioning attached. And a subsidiary or acquired brand gets collapsed into its parent, so answers about your product name return the parent company’s description.
All four are detectable in an afternoon and none are detectable by looking at your own website. You find them by asking the engines directly and reading what comes back, which is the diagnostic described later in this section.
What has to be consistent, and how consistent is consistent?
Every surface you control should carry the same name string, the same one-line description, the same category and the same core facts, character for character where it is reasonable to be that strict. This sounds like pedantry and is not: an entity is assembled by matching strings across sources, and “Rowan Freight Audit Ltd” on the invoice, “Rowan Freight” on LinkedIn and “RowanFreight” on the industry directory are three strings that an engine has no obligation to unify.
The core set is small enough to fit in a document and should live in one: legal name, trading name, the one-line description, the category words you want associated with you, the founding date, the head-office location, the founders and named experts with their exact role titles, the product and service names, and the canonical logo file. Write that document once, and treat it as the source of truth every profile, byline, directory listing and press mention is copied from. The failure mode is not that anyone disagrees with the facts; it is that six people wrote six profiles over four years and nobody was working from the same sheet.
How do you get third parties to corroborate you — and what will we not do?
Corroboration comes from surfaces you do not control, which is exactly why it counts. The routes that work are ordinary and slow: accurate profiles on the industry directories your buyers actually use, complete and current professional profiles for your named people, membership listings for associations you genuinely belong to, partner and integration pages on the sites of companies you genuinely work with, conference and podcast appearances that produce indexable transcripts and show notes, and earned coverage in publications that cover your category. Each of these repeats your core facts in somebody else’s voice on somebody else’s domain, which is the property that makes it useful.
Wikipedia and Wikidata sit in this list for genuinely eligible organizations and nowhere near it for everyone else. Wikipedia requires notability demonstrated by significant coverage in independent, reliable sources, and has an explicit conflict-of-interest policy governing anyone writing about a business they are paid by. Wikidata has its own notability policy requiring a serious, publicly available reference or a clear structural need. Attempting to place an entry for a business that does not meet those bars is against the rules of both projects and does not work: entries get reverted or deleted, often within days, and the attempt can leave a paid-editing flag attached to the brand that is considerably worse than absence. We do not create, edit, or commission edits to Wikipedia or Wikidata entries for clients. If a client is genuinely notable, the correct sequence is that independent coverage exists first and an independent editor writes the entry because the subject warrants one. If a client is not notable, the honest answer is that this route is closed and the budget belongs elsewhere.
The same reasoning rules out the rest of the category’s shortcuts. Fake or incentivized reviews, manufactured forum and Reddit accounts, paid placements dressed as independent coverage, and bulk directory submissions all attack the corroboration mechanism by simulating agreement rather than earning it. They are deceptive, platforms remove them, and they are self-defeating: a sudden cluster of similar mentions is the exact signature that both moderation systems and human fact-checkers look for. We decline this work, and if you are being offered it, the offer tells you something about who is making it.
How does structured data assert an entity?
Structured data is where you state your entity facts in a form a machine reads without inference, and it is the cheapest high-confidence signal available. Organization on the homepage carries name, legalName, alternateName, url, logo, description, foundingDate, address and sameAs. Person markup on team and author pages ties your named experts to the organization and to their own profiles. Product and Service name what you actually sell, in the words you want repeated. The connective tissue is sameAs: an array of URLs pointing at your profiles elsewhere, which is you telling the engine explicitly which external identities are the same entity as this one.
Three implementation rules prevent most of the waste. The markup must be in the served HTML, not injected by a tag manager after load, because a crawler that does not execute JavaScript will not see it. Every sameAs URL must return 200 and must actually be your profile — a dead link to a service you stopped using is an assertion that is now false. And the values must match the source-of-truth document exactly; structured data that contradicts the visible page is worse than no structured data, because you have now asserted two different entities from one URL.
How do you detect and fix a confused engine?
You detect entity confusion by asking the engines and recording the answers verbatim with the date, because there is no dashboard for this and no signal on your own site. Run the same four prompt shapes across ChatGPT, Perplexity, Google AI Mode and one other: “What is brand?”, “Who founded brand?”, “What does brand do?” and “Is brand the same company as similar name?”. Read the answers for wrong facts, merged companies, retired names, wrong locations and products you discontinued. Repeat monthly, keep the transcripts, and date them — a description that was right in June and wrong in September tells you something happened that you need to find.
The remedy is a set of moves, none of them fast. Publish an explicit disambiguation sentence on your About page naming the confusion directly — a plain statement that your company is not affiliated with the similarly named one, with enough distinguishing detail (category, location, founding year) to separate the two. Add alternateName for former names and known variants so the old string resolves to the current entity. Tighten sameAs so the profile set is unambiguous and every link is live. Get the correct facts restated on independent surfaces, because the corpus is what has to change and your own page is one vote in it. Then wait: corpora update on their own schedule, browsed answers refresh faster than model recall, and anyone promising a fixed timeline for a corrected description is promising something they cannot control.
| Entity signal | Where it lives | How to verify it is correct today |
|---|---|---|
| Legal and trading name | Company registry, About page, invoices, Organization.legalName and name | Put all four strings side by side in one document. A Ltd-versus-Limited mismatch is enough to split an entity. |
| One-line description and category | Homepage <title> and meta description, Organization.description, every third-party profile | Paste each version into one document and read them together. They should be paraphrases of one business, not descriptions of several. |
| Structured data | Organization, Person, Product, Service JSON-LD in the served HTML | Fetch the raw HTML and confirm the JSON-LD is present without JavaScript; validate it; check every value against the source-of-truth document. |
sameAs profile links | Homepage JSON-LD, pointing at every external profile | Request each URL and confirm a 200 and that the page is genuinely yours. Remove links to services you have abandoned. |
| Named people and roles | Person markup, author bylines, team page, professional profiles, conference bios | Same name spelling and same role title everywhere. Check bylines specifically — they drift fastest and are rarely reviewed. |
| Location and address | PostalAddress, site footer, business listings, directory entries | One address format, one phone format. Compare the footer against every listing rather than against your memory of them. |
| Founding date and history | foundingDate, About page, registry record, press coverage | Confirm the registry date and the About page agree; if a rebrand or acquisition changed it, state both dates explicitly rather than picking one. |
| Products and services | Product and Service markup, services pages, third-party listings | List every product name in use across all surfaces. Retired names still live in old listings and keep resurfacing in answers. |
| Former names after a rebrand | alternateName, redirects from the old domain, an explicit “formerly” sentence on the About page | Search the old name and see what the engines return. If the old identity still answers, the bridge has not been built. |
| The engine’s own belief | ChatGPT, Perplexity, Google AI Mode and one other | Run the four disambiguation prompts monthly, record the answers verbatim, and date the transcript. This is the only direct read of what the engine thinks you are. |
Answer-ready content architecture
Answer-ready content architecture is the practice of writing so that a self-contained span can be lifted out of a page and used as an answer without the surrounding context. It is a different discipline from writing well, and occasionally in tension with it: the paragraph that builds elegantly toward its point is the paragraph whose point arrives in a chunk that no longer contains the setup. The unit of work is the passage, the test is whether a stranger reading sixty words cold understands them, and the failure mode is a page that is genuinely useful to a human and unusable to an extractor.
What does passage-level design actually require?
Six rules cover most of it, and they are mechanical enough to check against a draft.
- Answer the heading in the first sentence. Every section opens with a sentence that answers its own heading completely, before any qualification. If the heading asks how long something takes, the first sentence contains a duration.
- Name the subject in every sentence that carries a claim. Not “it typically takes four weeks” but “a freight invoice audit typically takes four weeks”. Repetition that reads slightly heavy to a human is what makes a span portable.
- No anaphora across an opening. Never begin a paragraph with “This means”, “That is why”, “It does this by” or “As mentioned above”. Each of those makes the first sentence of a chunk depend on a sentence the chunk does not contain.
- Put the definition before the context. Definitional sentences go first because extractors read top-down and stop early. The history, the caveats and the exceptions go after.
- Phrase headings as buyer questions. A heading is a retrieval anchor as well as a signpost. “Our approach” anchors nothing; “How is a freight invoice audit different from a rate negotiation?” anchors a query someone actually runs.
- Keep one claim per passage. A paragraph making four points gets chunked into fragments of four half-arguments. A paragraph making one point survives being cut anywhere.
Which formats get quoted most readily, and why?
Four formats get lifted far more readily than prose, and they share one property: each is already a self-contained unit before the chunker touches it. A definition is a complete assertion in one sentence. A comparison table is a set of facts with their own labels, so a row survives being extracted from its table. A stepwise procedure carries its own ordering, so step three still reads as step three. An FAQ block is a question with a bounded answer, which is the exact shape of what the engine is trying to produce.
| Format | Why it gets quoted | What makes it fail anyway |
|---|---|---|
| Definition | A single sentence that fully answers “what is X” is directly usable as an answer with no editing | Definitions that open with the brand rather than the concept; definitions that need the next sentence to complete them; three competing definitions on one site |
| Comparison table | Rows carry their own column labels, so each row survives extraction on its own; tables also lift cleanly as structure | Cells containing a single word with no context; a comparison against unnamed “other providers”; tables built as images or CSS grids rather than <table> markup |
| Stepwise procedure | Ordered steps carry their sequence with them and map directly onto “how do I” questions | Steps that reference earlier steps by pronoun; steps that are aspirations rather than actions; procedures with no stated starting condition |
| FAQ block | A question with a bounded answer is already the shape of the output the engine is assembling | Answers that open with “Great question” or restate the question before answering; answers written for keyword coverage rather than because anyone asks |
| Narrative prose | Carries argument, nuance and the reasons behind a recommendation — the part that makes a page worth trusting | Chunks badly by default. Use it, and make sure the load-bearing claims also exist somewhere in one of the four formats above |
Why is specificity the strongest citation driver?
Specificity drives citation because a citation is a pointer to a source, and only a specific claim needs a source. “Freight invoice audits typically recover meaningful savings” needs no attribution because it asserts nothing checkable. “Duplicate billing and dimensional-weight rounding account for the two largest error classes in parcel invoices we review” needs attribution, because a reader who wants to verify it has to go somewhere — and the somewhere is you.
The categories that do the work are named entities, dates, figures, exact settings and named methods. Write the version number, not “the latest release”. Write the menu path, not “in the settings”. Write the date the thing changed, not “recently”. Write the range and say what it depends on, rather than writing “varies”. Every one of these turns an adjective into a fact, and facts are the only thing a citation can anchor to. The corollary is uncomfortable and worth stating plainly: vague copy is not merely weaker in this channel, it is structurally uncitable, and a page made entirely of it will never be quoted no matter how much of it you publish.
Should you build a page per keyword variant?
No — answer engines resolve intent rather than match strings, so a set of near-duplicate pages targeting keyword variants is worth less than one page that fully answers the underlying question. The engine rewrites the buyer’s question into several queries before it retrieves anything, which means the variant you built a page for is frequently not the query that runs. What the retrieval step then finds, in the case of a variant-page site, is four thin pages that each answer a quarter of the question and compete with each other for the same slot.
The consolidation test is whether two pages would have the same first sentence if both were written to answer their own heading properly. If they would, they are one page. Merge them, redirect the losers, and spend the recovered effort on depth — the edge cases, the failure modes, the procedure, the table of variants — because depth is what produces the specific spans worth citing. One page that answers a question completely can be cited for every phrasing of that question; four pages that each answer it partially can be cited for none of them.
What does a stale page cost?
A stale page on a moving subject costs more than no page, because it competes for the slot and then supplies an answer that is wrong. When an engine summarizes something volatile — pricing, platform features, regulation, availability — it prefers sources that look current, and a page that confidently describes last year’s state of the world is a liability with your name attached: if it does get cited, the engine attributes an incorrect statement to your brand in front of a buyer.
Maintenance cadence should follow volatility rather than a uniform calendar. Pages describing a platform that ships changes monthly need review monthly. Definitional and methodological pages can go a year. The mechanism is a review date per page, an owner, and a short changelog entry recording what actually changed — and a visible last-updated date that reflects a real edit. A timestamp bumped on unchanged text is detectable by anything that diffs content between crawls, and it trades a small short-term signal for the risk of being treated as a source that misrepresents its own currency.
A worked before and after
The example below is illustrative and the business is invented. “Rowan Freight Audit” is not a client and is not a real company; the freight specifics in the rewrite are placeholders standing in for whatever your own verifiable facts are. What is real is the set of edits, which is the same set we make on client pages.
Before — a generic marketing paragraph. This is the shape most service-page copy arrives in:
Before (illustrative)
At Rowan Freight Audit, we’re passionate about helping businesses take control of their shipping spend. Our industry-leading solutions draw on decades of combined expertise to deliver seamless savings across your logistics operation. We work with clients of all sizes to unlock the full potential of their freight budget. Whatever your challenges, our dedicated team is here to help. Get in touch today to find out how much you could be saving.
Five sentences, ninety words, and nothing an engine could quote. There is no definition, no subject in most sentences, no figure, no date, no named method, no scope, and no claim that could be checked or would need a source. An engine that retrieves this page learns that a company exists and moves on to a page that says something.
After — the same paragraph rewritten as a citable passage:
After (illustrative)
A freight invoice audit is a line-by-line review of carrier invoices against your contracted rates, accessorial schedules and service guarantees, run to identify billing errors and file refund claims. Rowan Freight Audit reviews parcel and LTL invoices for distributors shipping between 5,000 and 50,000 parcels a month. Four error classes account for most of what Rowan recovers: late-delivery service failures, residential-surcharge misclassification, dimensional-weight rounding, and duplicate billing across consolidated invoices. Refund claim windows are contractual rather than universal — check the service-guarantee clause in your own carrier agreement before assuming a deadline. Rowan works on a contingency basis and does not audit air freight or international customs charges.
Six changes did that work, and each one is portable to any page.
- The definition moved to the front. The passage opens by defining the service, in a sentence that is complete on its own. That sentence alone can answer “what is a freight invoice audit” with no other context.
- The subject is named, repeatedly. “A freight invoice audit”, “Rowan Freight Audit”, “Rowan”. No sentence opens with a pronoun pointing backwards, so no chunk boundary breaks it.
- Adjectives became figures. “Clients of all sizes” became a stated range of 5,000 to 50,000 parcels a month, which is both a qualification signal for the buyer and a checkable fact for the engine.
- The vague benefit became a named list. “Seamless savings” became four named error classes. A list of named things is the single highest-yield edit available, because each item is independently quotable.
- A hedge became a specific one. Rather than dropping the caveat or writing “terms may vary”, the passage says the window is contractual, says where to look for it, and declines to state a number it cannot stand behind.
- Scope was stated, including the exclusions. Naming what Rowan does not audit disqualifies the wrong buyer and gives the engine a boundary — which is exactly the kind of detail that gets quoted when someone asks whether a service covers their case.
The rewritten passage is 110 words against the original’s 90, so this is not a length exercise. It is a density exercise: the same space carrying facts instead of adjectives. Run the test on your own copy by deleting every sentence that contains no name, number, date or named method, and reading what is left.
Why do we publish the whole method instead of gating it?
We publish the full method on pages like this one because a page that withholds its substance has nothing to quote. The logic is unavoidable once you accept how citation works: an engine cites a specific claim, a gated page contains no claims, and a page whose value sits behind a form is, to a retrieval system, an empty page with a call to action on it. Every ungated procedure, table and definition on this site is a candidate span; every gated PDF is not.
The commercial objection — that publishing the method lets readers do it themselves — is correct, and we accept it. Plenty of readers should run this themselves; the method is not the scarce thing. What is scarce is the labor of running it properly across a real site with real infrastructure, real legacy content and a real security team, and the judgement about which of thirty findings actually explains why you are absent from the answers you care about. If reading these five sections is enough for you to fix your own site, that is a good outcome and you should go and do it. Our guide to appearing in ChatGPT answers is the condensed version to start from.
Citation and authority engineering
Citation and authority engineering is the workstream that happens off your own website: getting independent sources to describe your company accurately, specifically and often enough that a model assembling an answer finds corroboration rather than only your own claims. It is the half of generative engine optimization that cannot be shipped by a developer, cannot be finished in a sprint, and cannot be bought outright — and it is the half that decides whether you are one of the three to five brands named in an answer or none of them. It is also the dirtiest corner of this category, because the shortcuts are obvious and the detection is invisible until it isn’t.
Why publishing on your own site is not enough
A language model answering a commercial question is doing something closer to weighing testimony than to ranking documents. Your own site tells it what you assert about yourself. Independent sources tell it whether anyone else agrees. When the question has stakes — someone is deciding what to buy — the corroborated claim outranks the self-made one, exactly as a buyer weights a peer’s account over a sales page. A claim that appears only on your domain is a claim with a sample size of one.
The published analysis of how brands surface in AI answers puts numbers on it. Roughly 85% of brand mentions originate on domains the brand does not own, and brands are around 6.5× more likely to be surfaced through third-party sources than through their own content (third-party analysis, attributed in our guide to why ChatGPT recommends your competitor). Two separate filters are running: a shortlist decision about which brands are safe to name, then a grounding decision about which passages support the claims. Your website participates mostly in the second. The first is settled elsewhere, before your content is ever consulted.
This is the most common strategic misdiagnosis in the category, and it is expensive because it looks like diligence. A brand publishes forty pages, restructures them for extraction, adds schema, and remains absent from every evaluation-stage answer — because the problem was never that its content was unciteable. The problem was that when the model looked for independent evidence that this company existed as a real option in the category, it found nothing. More content cannot fix a corroboration gap. The mechanism behind both filters is set out in how brands appear in ChatGPT responses.
A claim only you make carries the weight of an advertisement. A claim three independent sources repeat carries the weight of a fact.
Which surfaces actually count
Legitimate citation surfaces share one property: something real had to happen for you to appear there. That property is also what makes them slow, and it is why they are durable once earned. Six categories carry most of the weight in most B2B and considered-purchase categories.
- Independent comparison and roundup articles. The highest-value surface on commercial questions, because query fan-out decomposes “what should I use for X” into sub-queries that look almost exactly like comparison headlines. Earned by approaching publishers whose pieces your own measurement shows are actually being cited, with a specific and honest pitch about who your product is genuinely best for. Many maintain these as living documents. The pitch that works names the segment you lose; the pitch that fails claims you win everything.
- Trade and industry publications, through genuine expert contribution. A named practitioner writing something only a practitioner could write. This is a slow relationship business and it does not scale, which is precisely why it corroborates. A contributed article that could have been written by anyone in the category adds a byline and no authority.
- Original research and data other people want to cite. The single content activity that reliably converts publishing budget into third-party citation, because other people need something to cite and unique numbers are the most citable object on the internet. A benchmark, a survey of your own customer base, an aggregate figure that exists nowhere else. Publish the methodology alongside it or it will not be trusted or reused.
- Conference talks and podcast appearances. These produce transcripts, show notes, and secondary write-ups — text that sits on domains you do not own and describes you in someone else’s words. Earned by having something to say, which is a real constraint and not a euphemism.
- Credible directories and category listings. The narrow, curated ones where inclusion means something: a professional body, a marketplace with a review requirement, a category index maintained by a publication. Worth doing carefully once and then keeping accurate. Worth nothing in bulk.
- Documentation and integration listings with real partners. If your product integrates with something, the partner’s directory, changelog and docs are high-trust surfaces describing you in a functional, factual register. They are also frequently stale, because nobody on either side owns them. Auditing every partner listing you already have is one of the cheapest pieces of work in this discipline and it is skipped almost universally.
Notice what is missing from that list: guest-post networks, sponsored “top 10” placements sold by the article, and syndication farms. They are excluded because they are transactions dressed as evidence, and the corpus increasingly treats them that way.
What review platforms actually do — and it is not what most teams think
Review platforms are among the most heavily retrieved sources on commercial questions, for three structural reasons: they are explicitly evaluative, they are structured and comparable, and they are independent of the vendor. That combination is close to ideal for a system assembling a recommendation. But the operative output is not the star rating. It is the language.
When a model describes your product — “good for small teams, weaker on reporting, priced per seat” — that sentence is usually a synthesis of how reviewers described you, not a summary of your homepage. This is why review strategy is a description problem as much as a presence problem. Four properties matter, roughly in this order:
- Specificity. A review that names a use case, a team size, a workflow and an outcome supplies a groundable passage. “Great product, highly recommend” supplies nothing to any answer, ever. If your review request template asks for a rating and nothing else, you are generating volume with no citable substance.
- Recency. A cluster of reviews from this quarter outweighs a larger set from two years ago, particularly in fast-moving categories. Recency is also how you correct a stale description: new reviews describing what the product is now gradually outweigh old ones describing what it was.
- Volume. Signals that the product is genuinely in use rather than newly launched. It is a threshold effect more than a linear one.
- Distribution. Presence across several platforms beats concentration on one, because different platforms are retrieved on different sub-queries.
The programme that works is unglamorous: a standing request at the point of demonstrated value — after a successful onboarding, a renewal, a support resolution — with a prompt that asks the reviewer what they use it for and what changed. Steady accumulation beats a burst, both because platforms scrutinise bursts and because recency keeps mattering after the campaign ends.
What we will not do, and why the practical reason matters as much as the ethical one
Four hard lines
We do not astroturf community platforms. No accounts created to recommend a client, no seeded threads, no undisclosed participation. Where a client is genuinely useful in a discussion, they participate transparently as the vendor, or not at all.
We do not generate, incentivise in bulk, or fabricate reviews. Review requests go to real customers after real usage, with no scripting of the content and no reward conditional on sentiment.
We do not attempt Wikipedia entries for businesses that are not notable under Wikipedia’s own criteria, and we do not edit articles about clients. Notability is decided by independent coverage that already exists; if it does not exist, the entry is not a task, it is a symptom.
We do not buy links or placements and present them as citations. Paid guest-post networks, link packages, and pay-to-be-included “best of” lists are advertising. If a client wants to buy media, that is a legitimate decision and a different budget line — it is not authority work and we will not report it as such.
The ethical case is straightforward and we will not belabour it. The practical case is the one buyers underweight, so it is worth stating precisely: platforms remove manipulated content at scale, and a brand described by a poisoned corpus is harder to fix than a brand that was simply absent.
Three mechanisms make that true. First, detection is retrospective and it is automated. Review platforms and community platforms run pattern detection across accounts, timing, phrasing and IP clustering, and they act in sweeps rather than case by case — which means a campaign can look successful for months and then be removed in a single afternoon, taking legitimate content with it. Second, the manipulated material is generic by construction. Fabricated reviews and seeded threads say nothing specific, because specificity requires having actually used the product, so even undetected they contribute almost nothing to the description a model synthesises. You are paying for volume in a system that weights substance. Third, and worst, removal is not a return to zero. A community that has removed astroturfed threads about your brand now has moderators and members who associate the brand with the incident, and that association is itself text on a domain you do not own. Absence is a blank page. A poisoned record is a page with the wrong thing on it, and models are considerably better at finding text than at knowing it should be discounted.
There is also a plain commercial argument. Corroboration built honestly is difficult for a competitor to erode, because they would have to displace real coverage with real coverage. Corroboration built by manipulation is fragile by design and it transfers no defensibility. You are renting a position you cannot keep.
When the sources that describe you are wrong or outdated
Being described inaccurately is a distinct failure from being absent, it is more common than teams expect, and it is the cheapest thing on this page to fix. The mechanism is entity instability: a model builds a representation of what your company is from descriptions across many sources, and when those disagree — because your positioning changed, your pricing moved, a product was discontinued, or an acquisition rewrote the org chart — it reaches for whichever description is best corroborated rather than whichever is current. Old, widely-repeated facts beat new, thinly-sourced ones.
The correction routes that are legitimate, in the order we run them:
- Write one canonical description and use it verbatim. Category, audience, differentiator, in plain language, on a page that states what you are, who it is for and what it costs. Every property you control then repeats it word for word. Variation across your own properties is a self-inflicted entity problem.
- Update the profiles you already own. Review platform vendor profiles, directory entries, partner listings, social bios, conference speaker pages. These are frequently years stale, they are heavily retrieved, and you can change them today without asking anyone.
- Ask publishers to correct factual errors. Most will, particularly on pricing, product names, funding and headcount. Send the correction with a source, not a request for better coverage. Editors handle factual corrections routinely and ignore repositioning requests, and conflating the two is how the relationship gets burned.
- Publish the correction yourself, dated. A changelog, a pricing page with an effective date, an announcement post. This gives correcting sources something to cite and gives the model a recent, unambiguous version of the fact.
- Re-check after re-crawl, not immediately. Descriptions update as sources are re-crawled and, for the conversational surface, only across model updates. Expect weeks for the retrieval-grounded surface and considerably longer for anything drawn from the model’s trained parameters.
What is not a legitimate correction route: contacting a review platform to have unfavourable but genuine reviews removed, pressuring a publisher over an accurate negative assessment, or attempting to suppress a community thread. If the description is unflattering and true, the fix is upstream in the product or the customer experience, and we will say so rather than take the work.
| Surface | How a citation is legitimately earned | Typical time to effect | Risk if done badly |
|---|---|---|---|
| Independent comparison and roundup articles | Approach the publishers your own citation tracking shows are actually being retrieved, with a specific case for which buyer segment you are the right answer for | Weeks to secure inclusion, then weeks more before the updated page is re-crawled and cited | Paying for inclusion turns it into advertising the corpus discounts; over-claiming gets you edited out or described dismissively |
| Trade and industry publications | Genuine contributed expertise from a named practitioner, pitched to an editor who has a gap you can fill | 1–3 months from pitch to publication, longer for print-led titles | Ghostwritten filler builds a byline and no authority; paid placement networks attach your brand to a pattern publishers are actively pruning |
| Original research and data | Publish something quantitative nobody else has, with the methodology attached, then tell the people who cover the category | 2–6 months for citations to accumulate; the asset then keeps earning for years | Thin or self-serving data gets ignored at best and picked apart at worst; an unmethodical number that is later disputed damages the entity you are trying to stabilise |
| Review platforms | A standing request to real customers at the point of demonstrated value, asking for use case and outcome rather than a rating | 2–4 months of steady accumulation before weighting shifts | Incentivised or fabricated reviews breach platform terms, are removed in sweeps, and contribute nothing citable even before removal |
| Podcasts and conference talks | Having a specific, non-promotional thing to say and pitching it to the right show or programme committee | 1–4 months to appear; transcripts and write-ups index within weeks after | Pure pitch appearances produce transcripts that describe your marketing rather than your expertise, which is what then gets synthesised |
| Directories and category listings | Meet the inclusion criteria of the small number of curated indexes that matter in your category, then keep the entry accurate | Days to weeks, and near-immediate for correcting an existing stale entry | Bulk directory submission is a recognised low-quality pattern; stale entries actively misdescribe you and are heavily retrieved |
| Partner documentation and integration listings | Build or maintain the integration, then work with the partner to get the listing, docs and changelog right | Weeks, mostly gated by the partner’s release cycle rather than by anything you control | Listings for deprecated integrations describe a product you no longer sell, and neither side is monitoring them |
| Practitioner communities | Build something people discuss unprompted; participate transparently as the vendor only where you add real information | The slowest surface here — typically 6 months or more, and largely outside your control | Astroturfing is recognisable, removed, and leaves a moderator-visible record that is worse than never having appeared |
What this means for your programme
Budget for citation work the way you would budget for PR and customer research, not the way you would budget for content production. The unit of work is a relationship, a dataset or a real customer outcome — none of which can be produced on a content calendar. If a supplier’s citation plan consists entirely of things they can execute alone from a laptop, they are selling you the part of the problem that was never the constraint.
Designing the prompt set
The prompt set is the measurement instrument for the entire engagement, and choosing it badly makes every number that follows decoration. A prompt set is the fixed list of questions run repeatedly against each assistant to establish whether your brand is named, how it is described, which competitors appear beside it and which sources the answer was grounded in. It is the closest thing this discipline has to a keyword list, and the resemblance is superficial enough to be dangerous: keyword volume is measured externally and independently of you, whereas a prompt set is authored, which means it can be authored to flatter. This is the most under-specified part of most GEO proposals, so it is the part we specify hardest.
Where the prompts come from
Prompts are derived from how buyers actually ask, not from a keyword export. The two produce different artefacts. A keyword is a compressed fragment typed into a search box because search boxes reward compression. A prompt is a sentence typed into a conversation because conversations reward context, and it is consequently long, constrained and often embarrassingly specific: “we’re a 12-person dental practice and the front desk is drowning in appointment calls — what should we look at?” Running keyword-shaped prompts produces keyword-shaped answers that no real user ever sees, and the resulting visibility number describes a surface nobody is looking at.
Four sources produce prompts worth tracking, and all four are inside your business already:
- Real buying questions. Sales call recordings and transcripts, the first two minutes of discovery calls, inbound form free-text fields, and the questions support answers in the first week of a new account. Take the wording verbatim before anyone tidies it. The awkward phrasing is the point.
- The stages of an evaluation. A buyer moves from “I have a problem I cannot name” to “what category solves this” to “which options exist” to “which two should I shortlist” to “is this one a mistake”. Each stage produces a structurally different prompt, and a brand can be strong at one stage and invisible at the next. Our guide to visibility during buyer evaluation covers the stage where the revenue is decided.
- Competitor comparisons. Both directions. “X vs Y”, “alternatives to X”, and “best X for [your niche]” — including the versions where the competitor is the subject and you would have to be named as the alternative. These are the prompts where third-party weighting is harshest on vendor content and therefore where the corroboration gap shows up first.
- The objections that decide deals. “Is X worth the price”, “what are the downsides of X”, “why do people leave X”, “is X secure enough for a regulated industry”. Your sales team can list these from memory. They are the highest-stakes prompts on the panel because the answer is assembled almost entirely from independent sources, and because a bad answer here kills a deal silently.
Coverage across the funnel, and why branded prompts flatter you
A prompt set weighted toward branded prompts produces high numbers and teaches you nothing. If someone types your company name, the model has already been handed the entity; naming you back is a low bar and clearing it tells you almost nothing about whether you would be discovered by someone who has never heard of you. Branded prompts belong on the panel — they are how you detect a wrong description, which matters enormously — but as a small minority.
The weighting we start from, consistent with the panel design in our guide to tracking ChatGPT visibility: roughly 40% evaluation-stage prompts, 30% problem-stage, 20% category-generic, 10% brand-specific. The rule that keeps it honest is deliberately including prompts you expect to lose. A panel built from questions you already win measures your own comfort. The zero-coverage entries — prompts where competitors appear consistently and you never do — are the single most actionable output the whole instrument produces, and a flattering panel produces none of them.
Report the stages separately and never only in aggregate. Informational and commercial prompts behave differently enough that a healthy blended figure routinely conceals total absence at the evaluation stage, which is the only stage where a deal is won or lost.
How large should the set be, and how stable?
Two properties are in direct tension, and pretending otherwise is how suppliers end up reporting movement that is really instrument drift.
The set must be large enough to be representative. With fifteen prompts, one prompt is nearly seven percentage points of your citation rate, so a single answer changing its mind moves the headline number more than a quarter of real work would. Small panels also cannot be segmented: split fifteen prompts across four funnel stages and each stage has three or four, which is not a measurement, it is an anecdote with a percentage sign.
The set must be stable enough that week-over-week movement means something. Every prompt added or reworded changes the instrument, and a number produced by a changed instrument cannot be compared with last month’s. The temptation to keep adding is constant and mostly well-intentioned — a new competitor appears, a product launches, sales hears a new objection — and it is the most common way a visibility programme quietly loses its baseline.
How we resolve it in practice: start at 30 to 50 prompts, which is enough to segment into four stages and still have eight to twenty in each, and freeze the panel for a full quarter. Changes are versioned rather than edited — the panel gets a version number, the old and new panels run in parallel for one cycle, and both are reported during the overlap so the discontinuity is visible instead of hidden. Panel review is a scheduled quarterly event, not a running edit. Larger panels are better and cost more, in tooling and in review time; a manual programme with no tooling can be defensible at 10 to 25 prompts if you accept coarser resolution and stop reading small movements.
Each prompt is also run several times per cycle rather than once. Three to five runs per prompt per measurement cycle is our floor, and presence is reported as a rate across those runs rather than as a yes or no. A single run is a coin toss you have chosen to believe.
The same prompt behaves differently on each assistant
A single blended visibility number across ChatGPT, Gemini, Claude and Perplexity hides more than it shows, because the four systems retrieve differently, cite differently and weight sources differently. One will lean on review platforms for a commercial question where another reaches for community discussion; one surfaces links prominently and another summarises without them. A brand can hold a strong position on one assistant and be entirely absent on another, and a blended average will report that as mediocre-but-fine on all four, which is the one description that is true of none of them.
There is a second split inside ChatGPT specifically, and it matters for what you do about a bad reading. Conversational answers draw on the model’s trained parameters, so your presence there reflects how prominently and consistently you were described across the open web at training time — it moves over months and no publishing this quarter changes it this quarter. Browsing and search-grounded answers trigger live retrieval and respond to new material within days. Weak conversational presence is an entity problem; weak retrieval presence is a page problem. A tracker that blends the two into one score will have you applying the wrong remedy for a quarter.
So we report per engine, per stage, and we say which engine and which model version produced each reading. That produces more numbers than a single score, and it is the difference between reporting and diagnosis.
Volatility, and what it does to how often you should measure
Generative answers are not deterministic, and the instability is larger than most buyers assume. Third-party analysis in this category reports that only around 30% of brands stay visible across consecutive runs of the same prompt. The same order of magnitude shows up across major model updates, as covered in our guide to ChatGPT visibility metrics. Treat that figure as the headline risk in every reading you are shown.
Three consequences follow directly, and they are the reason our reporting looks the way it does.
- A single reading deserves almost no weight. If roughly seven brands in ten can appear in one run and vanish in the next, then one run tells you very little about your position and a great deal about the sampling. This is why presence is reported as a rate over repeated runs and why we will not send a screenshot of one good answer as evidence of anything.
- Measure weekly, decide monthly, judge quarterly. Weekly sampling is frequent enough to build a trend line without drowning it in variance. Daily measurement is almost entirely noise, because the underlying signals — third-party mentions, review depth, corroboration, entity consistency — move on a timescale of months. A supplier offering a daily visibility dashboard is offering you a more expensive random number generator.
- Movement smaller than the noise floor is not movement. Establish what your panel’s run-to-run variance actually is during the first cycles, then refuse to read changes inside it. A four-point rise on a 40-prompt panel with three runs per prompt may be a handful of coin flips. Say so, out loud, in the readout.
One more discipline the volatility forces: record which model version produced each reading. A drop that coincides exactly with a model release is a different diagnosis, with a different response, from a drop caused by a competitor publishing something. Without the version stamp, the two are indistinguishable and you will spend a quarter fixing the wrong thing.
| Prompt type | What it tells you | What it cannot tell you | How often to run it |
|---|---|---|---|
| Problem-stage (“how do I stop X happening”) | Whether you are present before the buyer even knows the category exists — the earliest and cheapest place to enter a consideration set | Nothing about commercial preference. Being named here is compatible with total absence when the buyer is choosing | Weekly, reported as its own segment |
| Category-generic (“what tools do teams use for X”) | Whether the model considers you a member of the category at all. The clearest read on entity-level standing | Which buyer you are right for, or how you compare. It is a membership test, not a preference test | Weekly |
| Evaluation-stage (“best X for a 20-person agency”) | Shortlist membership on the constrained questions where deals are actually decided. The highest-value segment on the panel | Why you were excluded. That answer is in the cited-source list, not in the presence number | Weekly, with the cited sources captured every run |
| Head-to-head comparison (“X vs Y”) | How you are positioned against a named rival, and whether the trade-off the model states matches the one you would state | Whether the comparison is being asked by real buyers at volume. Comparison prompts are easy to invent and easy to over-weight | Weekly for your top two or three rivals; monthly for the long tail |
| Objection and risk (“downsides of X”, “why do people leave X”) | What the corpus says about you when the question invites criticism. The fastest way to find a poisoned or stale description | How often buyers actually ask it. Low frequency, high consequence — read it qualitatively, not as a rate | Monthly, read in full rather than scored |
| Brand-specific (“is X any good”, “what does X do”) | Description accuracy: category, pricing, features, who it is for. The cheapest problem to find and usually the cheapest to fix | Anything about discovery. The entity was handed to the model in the question | Monthly, and immediately after any pricing or positioning change |
| Long-tail constrained (category + industry + size + budget) | Whether you own a defensible niche. Narrow prompts are frequently winnable when broad ones are not | Scale. A win on a prompt nobody asks is a win with no revenue attached to it | Monthly, and reviewed quarterly for whether the constraint still reflects real demand |
What the reporting actually contains
The reporting on a generative engine optimization engagement tracks four things separately and never blends them into a single score: citation frequency, share of voice against a named competitor set, which source URLs get cited, and sentiment or framing accuracy. Each answers a different question, each fails differently, and a composite “visibility score” that mixes them destroys the information that would have told you what to do next. Composite scores are also vendor-defined and non-portable — a 62 from one methodology and a 41 from another can describe an identical reality — which is why we report the components and let the trend do the summarising.
The four numbers, and why they must stay separate
- Citation frequency, or presence rate. The proportion of tracked prompt-runs in which your brand is named. This is consideration-set membership, and because answers name only three to five brands, it is close to binary per answer — there is no position four quietly collecting residual attention. Reported per engine and per funnel stage.
- Share of voice against a named competitor set. Your mentions divided by all brand mentions across the panel. This is the number executives ask for, and it carries one methodological trap worth stating on every report: share of voice requires a fixed competitor set. Adding a competitor to the list mechanically lowers your share without anything changing in the world, so the set is frozen alongside the panel and changes are versioned the same way.
- Which source URLs get cited. The most decision-useful output in the whole report and the one most often missing. Presence tells you that you are in the room; the cited-source list tells you which domains the model reached for to justify the answer. When a prompt shows zero coverage, the cited sources for that prompt are the work list — and if they are three review platforms, a community thread and an independent comparison, then publishing another landing page will not change the outcome.
- Sentiment and framing accuracy. Not a mood score. A record of what the answer said you are, who it said you are for, what it said you cost, and what it said your weakness is — checked against reality.
Being mentioned is not enough if the description is wrong
A brand described inaccurately in an AI answer is in a worse position than a brand that was not mentioned at all, and almost no reporting in this category measures it. The reason is mechanical: absence costs you a chance to be considered, while a wrong description actively disqualifies you in front of a buyer who will never tell you it happened. If the answer says you are enterprise-only and the buyer runs a ten-person team, they stop reading. If it quotes a price you abandoned eighteen months ago, you are eliminated on budget. If it attributes a limitation that belonged to a competitor, or to your own product two versions ago, you have been argued against by a source the buyer trusts more than your website.
The failure is invisible to a mention count, which is exactly why it survives. A dashboard reporting 61% presence looks like progress whether the 61% describes you correctly or describes a company you stopped being. We have seen the pattern often enough to check it first: a brand celebrating rising visibility while the answers name it as a point of comparison — the cheaper, thinner option a buyer should consider and then reject. Being named as the foil is a mention. It is also a loss.
So framing accuracy is graded on four attributes every cycle, from the raw answer text rather than from a score: category (are you described as the kind of thing you are), audience (is the buyer it recommends you to the buyer you sell to), facts (price, features, integrations, availability, ownership), and role in the answer (are you the recommendation, an alternative, or the cautionary example). This requires reading answers, not reading dashboards, and it is the single most common reason we ask for raw answer export from any tracking layer before we agree to use it. If you cannot read the actual text, you cannot audit the description, and the number you are being shown is unfalsifiable.
Business impact, and the honest problem with referral volume
Visibility is not revenue, and a report that stops at citation rate is a vanity report. Four business-impact readings are worth carrying, in ascending order of how hard they are to argue with:
- Referral traffic from AI surfaces, segmented into its own channel group rather than left to scatter across Referral and Direct. Judge it on behaviour — pages per session, conversion rate — rather than on volume.
- Branded-search movement. Rising AI presence typically shows up in branded query impressions before it shows up in referrals, because a large share of people who see your name in an answer go and look you up rather than clicking. This is the best leading indicator available and it costs nothing.
- Assisted conversions and self-reported attribution. A required “how did you hear about us” field on your form frequently captures more of this channel than any referrer ever will. It is biased toward what the buyer remembers, and it is still the only instrument that sees the majority who never click.
- Lead quality. Whether the inbound that arrives after a visibility improvement closes at a comparable rate. A visibility win that produces worse-fit leads is a targeting problem in your prompt set, and it will only be visible here.
Now the honest part. Third-party analysis puts LLM referral traffic at roughly 0.1% of website traffic today. That is a thousandth. Any supplier building a business case primarily on referral volume from AI assistants is either not looking at the number or hoping you will not. Referral traffic is real, it is worth instrumenting properly, and it is a weak primary metric at current volumes — it measures only the small clicking minority and misses the majority who absorb the recommendation and act on it later through a channel that looks like direct or branded search.
Which means the leading indicators have to carry the argument for the first two quarters, and we would rather say that plainly at the proposal stage than be asked about it in month four. The chain we report against is presence and citation share on evaluation-stage prompts, then branded-search movement, then self-reported attribution volume, then pipeline. Each link is weaker evidence than the one after it and available sooner. If your organisation cannot make a decision on leading indicators, this channel is difficult to justify at its current traffic contribution, and that is a legitimate reason to defer the whole programme.
The tracking layer
Measurement runs through measure.co, our AI-visibility tracker, from $89 a month. It runs a defined prompt set across ChatGPT, Gemini, Claude and Perplexity and reports citation rate, share of voice against a named competitor set, competitor benchmarking, a prompt-coverage dashboard, weekly reports and alerts. It is a separate product with published pricing, it works whether or not you engage us for the strategy work, and clients keep it if the engagement ends.
The thing we will state plainly, because the category generally does not: measure.co samples, exactly like every other tool in this category. It runs prompts and records answers. It has no privileged access to OpenAI, Google, Anthropic or Perplexity, no view of query volumes inside those systems, and no ability to see how often real users ask any given question. Nobody selling AI-visibility tracking has that, whatever the pricing page implies, and any supplier claiming a direct data relationship with a model provider should be asked to put it in writing.
What sampling well is worth is real, and it is worth being precise about. The value is a consistent methodology applied to a frozen panel over time — the same prompts, the same competitor set, the same number of runs, the same engines, versioned when it changes. That consistency is what turns a volatile signal into a trend you can act on. It is a measurement discipline, not a data advantage, and a client who understands the difference reads the reports correctly for the rest of the engagement.
What a monthly readout contains
Weekly numbers accumulate; the monthly readout is where they are interpreted. Ours contains six things, in this order:
- The four core metrics, per engine and per funnel stage, against the previous month and against the engagement baseline, with the noise floor stated so that movement inside it is labelled as such rather than narrated.
- Zero-coverage gaps. Prompts where competitors appear consistently and you never do, with the cited sources for each. This is the backlog, and it is the part of the document that generates work.
- Description accuracy findings, quoted verbatim from the answers, with any factual error routed to a specific correction task and a source to correct.
- Competitor movement. Who gained share, on which prompts, and what they appear to have published or earned. The cited-source list usually answers this without guesswork.
- Business-impact readings — branded search, AI referral behaviour, self-reported attribution counts, and lead-quality commentary where the sales cycle has produced enough closed outcomes to say anything.
- The experiment log. Covered below, and the most valuable page in the document by a distance.
The experiment log
The experiment log records what changed, when it changed, and what moved afterwards. One row per intervention: the date, the specific action (“added to [publisher] comparison”, “corrected pricing on three directory profiles”, “published category benchmark”), the prompts or segments it was expected to affect, the expected lag, and then the observed reading at the point that lag expires.
It exists because this discipline has long, variable lags and a noisy signal, which is a combination that makes causal claims very easy to invent after the fact. Without a log written in advance, every rise gets attributed to the most recent activity and every fall to a model update, and a year later nobody can say which of thirty interventions did anything. With one, you accumulate something genuinely rare in this category: a record of what worked in your category, on your prompt set, at your scale. That record is the actual asset an engagement produces. The visibility numbers are how it was measured.
The log is also the honest instrument for admitting that something did not work. An intervention whose lag has expired with no movement gets marked as such and stays in the record. A supplier whose log contains only successes has been editing it.
| Metric | What it measures | What it cannot prove | How to read a change |
|---|---|---|---|
| Citation frequency (presence rate) | Proportion of tracked prompt-runs naming your brand, per engine and stage | That anyone read the answer, or that the mention was favourable. Presence is not preference | Only above the noise floor of your panel, and only in the same panel version. Check the stage split before the headline |
| Share of voice vs named competitors | Your mentions as a proportion of all brand mentions across the panel | Market share, or that a competitor lost anything. It is a closed system that always sums to 100% | Confirm the competitor set is unchanged first. Adding a rival lowers your share with nothing happening in the world |
| Cited source URLs / citation share | Which domains grounded the answer, and what proportion of cited URLs were yours | Why a source was chosen, or what it would take to displace it. It shows the outcome, not the mechanism | Read as a work list, not a score. A new domain appearing repeatedly is a target; your own URLs disappearing is an early warning |
| Sentiment and framing accuracy | Category, audience, factual claims and your role in the answer — recommendation, alternative, or cautionary example | How a specific buyer reacted. It captures what was said, not what it cost you | Read the quoted text, never a score. Any factual error is a task with a named source to correct, regardless of trend |
| Zero-coverage gaps | Prompts where competitors appear consistently and you never do | Whether the prompt has commercial volume behind it. Your panel authored the question | Treat as the backlog. A gap closing matters more than an aggregate rising, because it is attributable |
| AI referral traffic | Sessions arriving from assistant surfaces, and how they behave once they land | Total influence. Third-party analysis puts it near 0.1% of site traffic, so it sees a small minority | Judge on conversion rate and pages per session, not volume. Volume changes here are usually too small to be significant |
| Branded-search movement | Change in branded query impressions and clicks in your search console data | Causation. PR, paid media, seasonality and product launches all move it too | Best leading indicator available. Read against the experiment log and annotate every other campaign that overlapped |
| Assisted conversions and self-reported attribution | Whether buyers say an assistant was part of how they found you, and whether those deals close | Precise credit. Self-report is biased toward the last memorable touch and toward channels people can name | Read as a monthly count and a proportion, needing roughly 100 responses before the split between channels is stable |
How to evaluate a GEO supplier, including us
Evaluate a generative engine optimization supplier on eight specific questions, and require an answer that describes a mechanism rather than an outcome. This section is written as a buyer’s checklist rather than a pitch, and it is deliberately usable against us — the last part applies every test to this offer and names where it comes up short. The category is roughly two years old, the deliverables are unstandardised, and the buyer usually cannot tell a methodology from a vocabulary. Eight questions close most of that gap.
1. Are the deliverables identical to a standard SEO retainer?
Ask: “Show me a sample month of deliverables, and tell me which of those items you were not already selling in 2023.”
The rebrand test is the fastest disqualifier in this category, because the cheapest way to enter a new market is to relabel the old one. If the scope is four blog posts, a technical audit, some internal linking and a rank report with an AI column bolted on, you are buying an SEO retainer with a new cover page. The overlap is genuinely real — crawlability, structure and clean information architecture all matter for retrieval — which is what makes the substitution so easy to hide.
A good answer sounds like: a scope with items that did not exist in an SEO retainer — a versioned prompt panel, a competitor set, per-engine reporting, a cited-source work list, a description-accuracy audit, a correction programme for third-party sources, and a citation-earning workstream with named target surfaces. A good supplier will also tell you which parts are ordinary SEO and price them as such, rather than charging a premium for a sitemap because the invoice says GEO.
2. Is the work confined to your own website?
Ask: “What proportion of the hours go to work on domains I do not own, and which domains are they?”
A website-only programme is optimising the minority signal. With roughly 85% of brand mentions occurring on domains you do not own, a supplier who never touches third-party sources, reviews or communities has scoped out the majority of the problem — usually because that work is slow, relationship-dependent and hard to invoice against a fixed monthly output. It is the most common structural weakness in GEO proposals and it is easy to detect: ask for the off-site work list and see whether one exists.
A good answer sounds like: a named list of target surfaces derived from your own citation tracking rather than from a generic list — these four comparison publishers because they were cited on eleven of your evaluation prompts, these two review platforms because your footprint is thin and stale, these partner listings because they describe a product you discontinued. Along with an honest statement that securing third-party coverage is an attempt, not a deliverable, and that the timeline is outside the supplier’s control.
3. Are they guaranteeing ranking, citation or placement?
Ask: “What exactly are you guaranteeing, and what happens if it does not occur?”
No supplier controls what a model says. Not us, not anyone. There is no submission endpoint, no inclusion request, no paid path into an organic AI recommendation, and no configuration that compels an assistant to name a brand. Outputs vary run to run — recall that third-party analysis reports only around 30% of brands staying visible across consecutive runs of the same prompt — which means that even a supplier who genuinely improved your position cannot guarantee a specific answer on a specific day. A guarantee of citation, placement, a visibility score or “position one in AI Overviews” is therefore either a misunderstanding of the mechanism or a claim made knowing it cannot be kept.
A good answer sounds like: guarantees about inputs and process, never outputs. Number of prompt cycles run, the panel and competitor set frozen and versioned, a stated number of target surfaces approached, correction requests sent, a monthly readout delivered on a date, and a defined response if leading indicators have not moved by an agreed point. Notice that all of those are things the supplier can actually control.
4. Can they show which prompts trigger you and which URLs get cited?
Ask: “Show me a real client report with the prompt-level data and the cited-source list. Redact the client name.”
A supplier without prompt-level and source-level tracking is working blind and reporting confidently. If the deliverable is a single visibility score with no panel behind it, there is no instrument, and the number cannot be audited, reproduced or acted on. The cited-source list is the specific thing to insist on: without it, every zero-coverage gap gets the same generic remedy, because nobody knows which domains actually grounded the answer.
A good answer sounds like: here is the panel, here is its version history, here are the runs per prompt per cycle, here is presence per engine and per stage, and here are the domains cited on each prompt ranked by frequency. Plus a willingness to hand over raw answer text on request. A supplier who cannot export the answers is asking you to trust a summary of evidence you are not allowed to see.
5. Do they measure how you are described, or only whether you appear?
Ask: “How do you detect that an answer described us inaccurately, and what happened the last time it did?”
Mention counting is the default because it automates cleanly. Framing does not automate cleanly — somebody has to read the answers — so it is quietly dropped from most scopes, and the failure it misses is the expensive one. A brand named with a wrong price, a wrong category, a wrong audience, or as the option a buyer should reject is losing deals while the dashboard reports improvement.
A good answer sounds like: a described process, not a sentiment score — answers read every cycle, four attributes checked (category, audience, facts, role in the answer), errors quoted verbatim and routed to a specific correction with a named source. And a concrete example of a correction they have actually run, including how long it took to show up and what did not change.
6. Do they report business impact, or only visibility?
Ask: “Which business metric would tell us this is working, and when would we expect to see it move?”
A programme reported entirely in citation rate can run for a year without anyone asking whether it produced revenue. The honest complication is that direct measurement is genuinely weak here: with LLM referral traffic at roughly 0.1% of website traffic per third-party analysis, referral volume cannot carry the argument, and a supplier who leans on it either has not checked or is counting on you not to. That is a reason to use leading indicators carefully, not a reason to skip business reporting entirely.
A good answer sounds like: a stated chain from visibility to branded search to self-reported attribution to pipeline, with expected lags for each link, an admission of which links are weak, and a commitment to report the business numbers even when they are flat. A supplier who claims clean attribution from AI answers to revenue is describing a capability that does not currently exist.
7. Are they chasing citations while the fundamentals are broken?
Ask: “What did you find in the first audit that you would fix before doing any citation work at all?”
Sequencing failures are expensive because they are invisible. Pursuing third-party coverage while your review footprint is thin and two years stale, your assistant-facing crawl access is blocked, your entity description contradicts itself across six properties and your pricing page states a number you abandoned last year, means paying for the slow work while the fast work sits undone. The unglamorous fixes are cheaper, faster and frequently produce the earliest visible change — because correcting a representation is easier than building one.
A good answer sounds like: a first month that is mostly diagnosis and correction, with citation work sequenced behind it, and a supplier willing to say that some of the first month’s findings are ordinary hygiene rather than GEO. If everything in the proposal is exciting, nothing in it is the actual first job.
8. Do they treat all the engines as one thing?
Ask: “Show me last month’s numbers broken out by assistant, and tell me where we are strongest and weakest.”
ChatGPT, Gemini, Claude and Perplexity retrieve differently, cite differently and weight sources differently, and a supplier reporting one blended figure is averaging away the only information that would change what you do. The follow-up question is sharper: ask whether they separate ChatGPT’s conversational answers from its search-grounded ones. A supplier who does not know the difference will keep prescribing content fixes for entity problems.
A good answer sounds like: per-engine numbers with the model version recorded, a stated view on which assistants matter for your specific buyer rather than all four by default, and different tactics for each — because a surface that leans on review platforms needs a different work list from one that leans on community discussion.
The same eight questions, applied to us
It would be dishonest to publish that checklist without running it, so here is where this offer does badly, or does not apply.
- We cannot guarantee any of it either. Everything in question three applies to us without exception. We guarantee cycles run, panels frozen, surfaces approached, corrections sent and readouts delivered on a date. We do not guarantee that you will be named, and we will not sign a contract that says otherwise.
- We do not write your content at volume. This engagement is diagnosis, prioritisation, citation and correction work, and measurement. If what you actually need is four articles a month, a content agency will do it better and cheaper, and we will tell you so on the call.
- We do not control third-party publishers, review platforms or communities — and we will not manipulate them. That means part of the work is an attempt with a real failure rate. A supplier who promises placements on domains they do not own is promising something they can only deliver by paying for it or faking it.
- Our measurement samples, like everyone’s. measure.co runs prompts and records answers, from $89 a month, with no privileged access to any model provider. If someone offers you privileged data, ask for it in writing.
- We are a small team, and Tarun is the delivery. That is the reason the judgement is consistent and it is also the capacity limit. We take on a small number of these engagements and we will say when we are full rather than staffing around it.
- This is organic AI visibility, not advertising. If your question is how to buy placement inside ChatGPT, the answer is that you cannot — sponsored units sit beside answers, not inside them, and organic recommendation is earned. That is a different service and a different page.
Who this is not for, stated plainly: businesses with no customers yet, no reviews and no independent coverage; teams that need results inside one quarter; organisations that cannot act on leading indicators; and anyone whose real problem is that their product is described accurately and unfavourably. Each of those is addressed in the timeline below, because for several of them the honest recommendation is to wait.
How long this takes, honestly
Generative engine optimization produces its first genuinely readable movement somewhere between six weeks and three months for most brands, and its first defensible business-impact signal considerably later than that. Nothing moves in week one. Anyone quoting faster is either describing a correction to something that was already broken — which is real, and which we do quote in days — or is reading noise as progress. The rest of this section is the realistic sequence, what is genuinely fast, what is slow by nature, and who should not start at all.
Why nothing moves in week one
Four independent lags sit between doing the work and seeing the number change, and they compound rather than overlap.
- Crawl and re-index lag. A page you change today is not a page the retrieval layer has seen today. Corrections propagate as sources are re-crawled, which is days to weeks depending on the domain’s crawl frequency — and the third-party sources that matter most are usually crawled on their schedule, not yours.
- Publication lag on everything you do not own. A comparison publisher who agrees to add you does it in their next update. A contributed article has an editorial calendar. A partner corrects a listing in their next docs release. Weeks, routinely, and entirely outside your control.
- Corroboration lag. One new source describing you correctly does not outweigh several old ones describing you wrongly. Entity representations shift when the balance of sources shifts, not when the first correct source appears.
- Measurement lag. With run-to-run volatility at the level third-party analysis reports, a real improvement takes several cycles to separate from noise. You are not waiting for the effect; you are waiting for the evidence of the effect.
The first quarter therefore has a shape, and the shape is front-loaded with work whose payoff arrives late.
- Weeks 1–2 — instrument and diagnose. Build and freeze the prompt panel and competitor set, run the first cycles to establish a baseline and a noise floor, capture the cited-source lists, and read the answers for description accuracy. Simultaneously fix assistant crawl access and rendering. Output: a baseline and a ranked work list. Output that is not produced: any improvement.
- Weeks 2–5 — correction and consistency. Canonical description written and deployed verbatim across owned properties; stale directory, partner and review profiles updated; factual corrections requested from external publishers; canonical facts page published with dates. This is the cheapest work in the engagement and the most likely to produce the first visible change.
- Weeks 4–10 — earned surfaces begin. Approach the comparison publishers your own tracking identified. Start the standing review programme. Begin the original-research asset. All three have long lags, which is why they start early rather than after the on-site work is finished.
- Weeks 6–12 — first readable movement, if it comes. Description accuracy improves first, then presence on brand-specific and long-tail constrained prompts. Evaluation-stage presence typically has not moved yet, and saying so in month three is the test of whether a supplier is reporting honestly.
- Month 3 onward — compounding, or a hard conversation. Either the leading indicators are trending and the experiment log shows which interventions did it, or they are not and the diagnosis was wrong. Both are legitimate outcomes at this point. Only one of them is usually reported.
The compounding argument, without the inflation
Compounding is real here and it is routinely oversold, so it is worth stating exactly. The technical and entity work — crawl access, rendering, consistent descriptions, clean structured facts — produces very little visible movement on its own. Its function is to be a prerequisite: it makes the later work capable of landing. A model that cannot retrieve your pages cannot cite them however good the content is, and a model holding an unstable idea of what your company is will not reliably shortlist you however many publishers mention you. Fix those and nothing happens. Fix those, then earn corroboration, and the corroboration converts into presence at a much higher rate than it would have.
That is the whole claim. It is not that effort accumulates into exponential returns. It is that the order of operations matters, and that work done in the wrong order produces a fraction of its effect and gets abandoned as ineffective. The corollary is uncomfortable for suppliers: the first month of a well-run engagement looks like the least productive one, and it is the one that determines whether months four through nine work.
The genuinely encouraging part is durability. Corroboration built over months is also slow for a competitor to erode, because displacing it requires them to earn equivalent coverage rather than to outspend you in an auction. Positions in this channel are harder to win and harder to lose than search rankings. That cuts both ways — a competitor who started eighteen months ago is correspondingly hard for you to displace — and any assessment of your timeline has to start with where they are, not with where you are.
What is fast, and what is slow by nature
The distinction is not effort. It is whether you control the surface.
Fast, because you control it: assistant crawl access and rendering; structured data and a canonical facts page; correcting a wrong entity description across your own properties; updating stale directory, partner and review-platform profiles; restructuring an existing page so its answers are extractable. Days to a few weeks, and the description fixes frequently produce the earliest movement in the whole engagement precisely because correcting a representation is easier than building one.
Slow, because someone else decides: independent corroboration; review depth and recency; inclusion in comparisons you do not publish; being discussed by name in practitioner communities; category presence from a standing start. Months, with real failure rates, and no amount of budget compresses them meaningfully. You can increase the number of attempts. You cannot shorten the cycle.
Who moves faster, and who should wait
A brand in a well-covered category with an established corpus footprint moves considerably faster than a new brand nobody writes about, and for some businesses the honest answer is that GEO is premature. This is the single most useful thing on this page for a prospective buyer to internalise, because it determines whether an engagement is optimisation or wishful thinking.
If independent sources already discuss your category and mention you occasionally — you have reviews, a few comparison appearances, some coverage, partners who list you — then the work is correcting, consolidating and extending an existing footprint. That is tractable inside a quarter or two, because the corpus already contains you and the job is to fix and amplify what is there.
If nobody writes about you at all, there is nothing to synthesise. A model cannot corroborate a company that no independent source has described, and no amount of publishing on your own domain manufactures that corroboration — that is the direct consequence of the 85% figure. What you need first is customers, reviews and coverage. Those come from selling, shipping and being interesting, not from a visibility retainer.
Specifically, we will tell you to wait if:
- You are pre-launch or pre-revenue. There is no customer base to generate reviews and no usage to discuss. Come back when there is.
- You have fewer than a handful of real reviews across all platforms. Review depth is a months-long accumulation and it is upstream of most commercial-prompt visibility. Start the standing review programme now; start paying for strategy later.
- Your category does not exist in the corpus yet. If you have genuinely invented a category, buyers are not asking about it and the prompts have no volume behind them. Category creation is a PR and market-education problem, and visibility work follows it rather than substituting for it.
- You need results this quarter. The lags described above are structural. If the budget is contingent on a result in ninety days, this is the wrong channel and paid media is the honest alternative — which is a different engagement, and we would rather point you at it.
- The unflattering description is accurate. If answers describe you as unreliable and the reviews saying so are real, the fix is the product or the support function. We are not able to change what the corpus says by changing what we publish, and we would be taking your money to fail.
Turning work down is not a rhetorical device here. We would rather tell you to wait than sell a retainer against a corpus that has nothing in it yet.
| Work | Realistic time to visible effect | What would make it faster | What would make it slower |
|---|---|---|---|
| Assistant crawl access and rendering fixes | Days to 3 weeks, once re-crawled | A frequently-crawled domain; server-rendered content; a developer who can ship the same week | A low-authority domain crawled rarely; client-side rendering; a release train with a monthly cadence |
| Entity and description consistency across owned properties | 2–6 weeks | Few properties, one owner, a decisive answer to “what are we” | Many stale profiles nobody has credentials for; an unresolved positioning debate; a recent rebrand still half-deployed |
| Canonical facts page and structured data | 2–6 weeks | Publishing prices and specifics openly rather than gating them | Pricing you will not publish, which removes the fact the model most needs and most often gets wrong |
| Restructuring existing pages for extraction | 3–8 weeks | Pages that already rank and are already retrieved — you are improving selection, not retrieval | Pages nobody retrieves. Restructuring an uncited page changes nothing, because selection was never the failure |
| Original research or benchmark data | 2–6 months for citations to accumulate | Data nobody else can produce, published with methodology, and actively taken to people who cover the category | A survey with a thin sample, or numbers indistinguishable from what already exists |
| Inclusion in existing third-party comparisons | Weeks to secure, then weeks more to be retrieved and cited | A clear, honest case for which buyer you are best for; an existing relationship with the publisher | Pitching every publisher generically; a category where the cited comparisons are closed or paid-only |
| Review depth and recency | 2–4 months of steady accumulation | A large happy customer base and a request built into the product moment where value lands | A small customer base; a long renewal cycle; asking for ratings instead of asking what the product is used for |
| Practitioner community presence | 6 months or more, and never fully controllable | A product people genuinely discuss unprompted, and staff who participate transparently and usefully | A category nobody discusses socially; any past astroturfing, which makes legitimate participation harder |
| Category presence from a standing start | 6–12 months, and contingent on real market presence | Funding news, notable customers, a distinctive point of view someone wants to quote | No customers, no coverage, no reviews. At that point the constraint is the business, not the visibility programme |
The test we apply before quoting
Before we quote, we run your category through a short version of the panel described above and look at what the corpus already contains about you. If the answer is “almost nothing, anywhere, from anyone”, we will say the engagement is premature and tell you what to do for the next two quarters instead. That test disqualifies a meaningful share of enquiries, and it is the reason this page spends as long explaining who should wait as explaining what we do.
What a programme delivers
A generative engine optimization programme produces a measured baseline, a prioritised roadmap, implemented changes, and a monthly record of what was tried and what moved. The last item is the one that separates a real programme from a retainer: in a category this volatile, an undocumented change is indistinguishable from noise six weeks later.
A baseline and a prompt set. The measurement instrument, defined and documented before any work starts: which prompts, why those, which engines, and how often they run. Without a baseline agreed in advance, every later claim of improvement is unfalsifiable — and that is the category’s most common trick, not usually a deliberate one.
A citation and competitor gap analysis. Which sources the engines currently cite for your priority prompts, which competitors appear where you do not, and which of those citations sit on surfaces you could plausibly earn. Some will not be earnable, and the analysis says which.
A crawl, access and entity audit. Whether the engines can reach you at all, whether your structured data asserts a coherent entity, and whether the facts about your brand are consistent across the surfaces you control. These are the gating layers, and they are frequently where a programme finds its cheapest wins.
A prioritised roadmap. Sequenced by dependency rather than by effort, because content architecture built on a broken entity foundation is wasted work. Each item carries what it is expected to move and what evidence would show it worked.
Implemented changes. Technical and entity fixes, content restructured into a form that can actually be quoted, structured data, and the off-site work of earning independent corroboration. Implementation, not a document telling your team what to implement.
A monthly readout and experiment log. Citation frequency, share of voice against named competitors, which source URLs got cited, and sentiment accuracy — reported separately, never blended into a single score. Alongside it, a dated log of every change made, so that when something moves you can say what preceded it.
Four things we will not put in a proposal
A guaranteed citation, mention or ranking in any AI system — nobody controls model output. A single blended “AI visibility score” that hides which engine moved and why. Any tactic that manipulates a third-party platform, including community astroturfing, fabricated reviews and Wikipedia entries for businesses that do not meet notability rules. And a forecast of referral traffic, because at roughly 0.1% of website traffic today, any such number would be invention.
How the engagement works
Programmes are scoped and priced individually, because the work genuinely differs by category, by how much the wider web already says about you, and by how much of the fix is off your own site. A brand with deep third-party coverage and a broken entity foundation needs a few weeks of technical work. A brand nobody writes about needs a year of earning coverage, and pretending otherwise with a standard monthly figure would be dishonest packaging.
That is a deliberate departure from how the rest of this site prices. Every other engagement here carries a published number, because quote-gating a price is usually a choice about which customer you want. This one does not, and the reason is not procurement theatre — it is that a single figure would be either meaningless or misleading given how far the scope moves between clients.
- Scoping call. Thirty minutes. What you sell, who decides, and what you believe is happening in AI answers today. We will tell you on that call if we think the programme is premature.
- Baseline. A prompt set built and run against the engines that matter to you, producing a measured starting position rather than an impression of one.
- Diagnosis and roadmap. Crawl and access, entity coherence, content citability, and the citation gap — sequenced into a plan with expected effects stated in advance.
- Proposal. Scope, price and cadence, written against that specific roadmap. If the diagnosis shows the work is small, the proposal is small.
- Execution and monthly review. Implementation against the roadmap, with a monthly readout and the experiment log.
Measurement, if that is all you need
measure.co is our AI-visibility tracking product: citation rate and share of voice across ChatGPT, Gemini, Claude and Perplexity, competitor benchmarking, weekly reports, alerts and a prompt-coverage dashboard. It is a standalone product with a published price, it requires no programme, and plenty of businesses genuinely need nothing beyond it. It samples like every other tool in this category — the value is a consistent method tracked over time, not access to data nobody else has.
When a GEO programme is the wrong purchase
This work pays off when there is already a corpus for a model to synthesise and something specific is going wrong inside it. Several situations fail that test, and this category is prone enough to overselling that naming them matters.
Almost nobody writes about you yet
This is the big one. Answer engines reflect what the wider web says, and roughly 85% of brand mentions come from domains a brand does not own — per the AirOps analysis of 548,534 pages across 15,000 prompts, which also found brands are around 6.5× more likely to be surfaced through third parties than through their own site. If you are pre-revenue, pre-review and pre-coverage, there is very little for a model to work with, and no amount of on-site optimization manufactures a corpus. Get customers, reviews and coverage first. We will say this on the scoping call and we say it often.
Your fundamentals are broken
Chasing citations while your reviews are thin, your basic technical health is poor or your business listings disagree with each other is spending at the wrong layer. Those are cheaper to fix and they gate everything above them.
You need attributable revenue this quarter
Then buy ads. That is not a deflection — it is what the rest of this site does, and managed ChatGPT Ads produces a click, a conversion event and a cost per conversion inside a month. GEO is a leading-indicator investment whose payoff is slow and whose measurement is honestly imperfect. Both things can be worth doing; only one of them answers a quarterly revenue question.
You want a guarantee
There is no ranking API and no supplier controls model output. If a guarantee is a requirement, no honest supplier can meet it, and the ones who accept the brief are the ones you should worry about most.
You only wanted to know where you stand
Then buy the tracking and skip the programme. measure.co at $89 a month will tell you your citation rate and share of voice, and for a lot of businesses that is the entire job. We would rather sell you the $89 product than a programme you did not need.
What to demand from any supplier here, including us
Ask to see the prompt set and how it was chosen. Ask whether mentions, citations, source URLs and sentiment are tracked separately. Ask what they will do about sources they do not own. Ask what would count as failure. And ask them to state plainly what they cannot measure — a supplier who cannot answer that last one has not thought about the category honestly, whatever their deck says.
Sources and further reading
The figures on this page come from third-party analyses rather than from any platform, because no answer engine publishes visibility data. Each is attributed at the point of use, and where the category’s published numbers are soft we have said so rather than borrow their confidence.
- AirOps — the analysis of 548,534 pages across 15,000 prompts behind the roughly 15% citation rate among retrieved pages, the 85% of brand mentions arriving from domains you do not own, and the 6.5× third-party effect. The most useful published dataset in this category.
- Published synthesis across AI-visibility tooling — the basis for the roughly 30% run-to-run brand persistence figure and for the description of API-sampling versus UI-scraping methodologies. These are third-party syntheses rather than first-party platform data, and they should be read as directional.
- Third-party traffic analyses — the roughly 0.1% of website traffic currently arriving from LLM interfaces. Directional, and moving.
- Our own guide library, which covers this discipline in more depth than a service page can: how brands appear in ChatGPT, appearing in ChatGPT answers, why ChatGPT recommends competitors, tracking AI visibility, visibility metrics, and buyer-evaluation visibility.
- measure.co — our own tracking product, published at $89/month, with the same structural sampling limits as everything else in the category.
One caveat that applies to the whole page. Nobody outside the model providers can observe what real users are actually shown; every visibility number in this industry, ours included, is an estimate built from a chosen prompt set run against a sampled interface. That does not make the measurement worthless — a consistent method tracked over time is genuinely informative about direction. It does mean that anyone presenting these numbers with the confidence of an impressions report is misrepresenting how they were produced.
Generative engine optimization FAQ
What does a generative engine optimization agency actually do?
A GEO agency works to make a brand findable, parseable and citable by answer engines such as ChatGPT, Perplexity and Google AI Overviews. In practice that is five workstreams: ensuring the engines can crawl and access you, establishing an unambiguous entity foundation, restructuring content so passages can be lifted and quoted, earning independent third-party corroboration, and measuring against a defined prompt set over time.
What is the difference between GEO, AEO and SEO?
Commercially the labels are near-synonymous and buyers use all of them. Read precisely, SEO optimizes for ranking a page in a results list, AEO optimizes for being the answer to a direct question, and GEO optimizes for being retrieved and cited when a model synthesises an answer from several sources. The practical test of whether a supplier understands the difference is simple: if their GEO deliverable list is identical to their SEO deliverable list, it is SEO with a new label.
How much does a GEO agency cost?
We scope and price programmes individually, which is a deliberate departure from the published pricing everywhere else on this site. The reason is that the work genuinely differs by category and by how much the wider web already says about you — a brand with deep third-party coverage and a broken entity foundation needs weeks of technical work, while a brand nobody writes about needs a year of earning coverage. A single figure would be either meaningless or misleading. Our tracking product, measure.co, is published at $89/month and stands alone.
Can you guarantee my brand appears in ChatGPT?
No, and nobody can. There is no ranking API to query and no supplier controls what a model says in any given response. A guarantee in this category is either a misunderstanding of the mechanism or a deliberate misrepresentation, and it is the clearest single signal to walk away from a proposal.
How do you measure AI visibility?
Against a defined prompt set, run consistently across the engines that matter to you, tracking four things separately: citation frequency, share of voice against named competitors, which source URLs get cited, and sentiment or framing accuracy. The honest caveat is that every tool in this category samples rather than observes real user sessions, so these are estimates on a prompt set you chose. That makes them useful as a trend and misleading as an absolute.
Is AI visibility tracking accurate?
It is consistent rather than accurate, and the distinction matters. Approaches split between API sampling, which reflects API-layer responses rather than exact user-facing output, and UI scraping that simulates a user. Neither observes what real users are shown. Published synthesis in the category reports that only around 30% of brands remain visible from one run of the same prompt to the next, so any single reading is close to meaningless and only the trend carries information.
How long does GEO take to work?
Technical access, structured data and correcting a wrong entity description can move within weeks. Independent corroboration, review depth and being cited by sources you do not control are slow by nature and are measured in quarters. Nothing meaningful moves in week one, and a supplier promising otherwise is describing a different category of work.
Why does ChatGPT recommend my competitors instead of me?
Usually because the wider web says more about them than about you. Answer engines assemble shortlists from entity-level signals across many sources, and roughly 85% of brand mentions come from domains a brand does not own — per AirOps' analysis of 548,534 pages across 15,000 prompts, which also found brands are around 6.5× more likely to be surfaced through third parties than through their own site. Our guide on why ChatGPT recommends competitors covers the diagnosis in depth.
Do backlinks still matter for AI search?
Independent corroboration matters, and links are one form of it — but the mechanism is different from classic link equity. What is being rewarded is that credible independent sources say the same thing about you, which is why a substantive mention without a link can outperform a link from a page nobody reads. Paid link schemes dressed as citations are both ineffective here and a reason we decline that work.
What is an llms.txt file and do I need one?
It is a proposed convention for offering LLM-oriented guidance about a site's content. Adoption across engines is uneven and its actual effect on citation is not established, so we treat it as cheap to add and honest to describe as unproven. Any supplier presenting llms.txt as a significant lever is overselling a file that costs an hour to write.
Does schema markup help with AI search?
Structured data helps mainly as entity assertion — it states unambiguously what your organisation is, what it offers and what other profiles refer to the same entity. That is genuinely useful when an engine is trying to resolve who you are. It is not a switch that produces citations, and it will not compensate for a brand the wider corpus barely mentions.
How do I get my brand cited in ChatGPT?
Be reachable, be unambiguous, be quotable, and be corroborated. Concretely: allow the relevant crawlers, make your entity consistent everywhere you control, write passages that answer their own heading in the first sentence so a span can be lifted without context, and earn independent sources that say the same things about you. Our guide on appearing in ChatGPT answers is the fuller treatment.
Which AI platforms do you monitor?
ChatGPT, Gemini, Claude and Perplexity through measure.co, plus Google AI Overviews. Which of those matter depends entirely on your buyers, and the same prompt behaves differently across them — so a single blended visibility score hides more than it shows. We report per engine.
How do you choose which prompts to track?
From real buying questions rather than from a keyword list: the questions that arise at each stage of an evaluation, the competitor comparisons your buyers actually run, and the objections that decide deals. The set has to be large enough to be representative and stable enough that week-over-week movement reflects the world rather than a changed instrument. A set weighted toward branded prompts will flatter you and teach you nothing.
Isn't this just SEO with a new name?
It is if the deliverables are identical, which is why that is the first question worth asking any supplier. The genuinely distinct work is entity and knowledge-graph coherence, passage-level content architecture built for retrieval rather than ranking, citation engineering on surfaces you do not own, prompt-set design, and multi-engine monitoring. A traditional SEO retainer contains none of those, and a GEO proposal that also contains none of them is a rebrand.
Can a small business do GEO, or is it enterprise-only?
Small businesses can absolutely do it, but stage matters more than size. If almost nobody writes about you yet, there is very little for a model to synthesise and the honest advice is to get customers, reviews and coverage first. A small business with genuine review depth and trade coverage is a far better candidate than a larger one with none.
Why pay for GEO when AI traffic is only a fraction of my visitors?
It is the strongest rational objection in the category and it deserves a direct answer rather than a deflection. Third-party analyses put LLM referral traffic at roughly 0.1% of website traffic today, so if you need attributable revenue this quarter you should buy ads instead — and we will tell you so. The case for GEO is that citation influences buyers before it produces measurable clicks, which is a leading-indicator argument, and leading-indicator arguments deserve more scepticism than they usually get.
Do you use Reddit, Wikipedia or review platforms to build visibility?
We earn presence on third-party surfaces legitimately and we decline the manipulative versions outright: no community astroturfing, no fabricated reviews, no Wikipedia entries for businesses that do not meet notability rules, no paid link schemes dressed as citations. The practical reason matters as much as the ethical one — platforms remove manipulated content at scale, and a brand described by a poisoned corpus is much harder to fix than one that was simply absent.
What do I get if I just want to know where I stand?
Buy the tracking and skip the programme. measure.co is $89 a month, published, standalone, and covers citation rate and share of voice across ChatGPT, Gemini, Claude and Perplexity with competitor benchmarking and weekly reports. For a lot of businesses that is the entire job, and we would rather sell you that than a programme you do not need.
How is this different from your ChatGPT Ads services?
Ads buy presence today; this work earns it. Managed ChatGPT Ads produces a click, a conversion event and a cost per conversion inside a month, with all the measurement caveats that channel carries. GEO is organic, slower, and measured with sampled estimates rather than platform-reported conversions. Both can be worth doing, but only one of them answers a quarterly revenue question, and we will not pretend otherwise.
Want to know where you actually stand?
Thirty minutes with Tarun. We will look at what the engines currently say about you and tell you honestly whether a programme is worth running — and if your brand has too little corpus footprint for this work to pay off yet, he will tell you to come back after you have more customers and more coverage.
Book a scoping call