How answer engines actually choose what to cite
Start with what is observable, because the rest is guesswork. When you ask ChatGPT or Perplexity something that needs current information, the assistant issues search queries of its own, fetches a small set of pages, and writes its answer from the text it just retrieved. The citations you see belong to that retrieved set, not to the whole web. You can verify it yourself: ask the same question twice, watch the sources shift, and notice how often the cited pages already rank for the query the assistant ran.
That mechanism splits the work into two separate problems. The first is getting retrieved: your page has to be findable by the query the model formulates, which is usually a cleaner and more literal version of the user question than anything a human would type. This remains a search problem, solved with indexation, relevance and authority. The second is getting quoted once retrieved: of the eight or ten pages fetched, only three or four typically end up cited, and the survivors are those whose text can be lifted into an answer with almost no rewriting.
What nobody outside the labs can tell you is the weighting. We do not know how much a system relies on training memory versus live retrieval, how citation slots get allocated, or how much this differs between ChatGPT, Perplexity, Gemini and the AI overviews inside Google. Anyone selling a reverse-engineered formula is describing a correlation at best. Act on the signals that are visible and stable across systems, and treat the rest as hypotheses you test. Our comparison of SEO and GEO covers the strategy; this article stays on execution.
Make your entity unambiguous
Before a system can recommend you, it has to be certain who you are. Entity resolution sounds abstract, but it fails in mundane ways: a company that trades under one name, invoices under another, and runs its product blog on a third domain hands a retrieval system three weak signals instead of one. Pick a canonical name, capitalisation included, and use it identically everywhere: your website, your social profiles, your press materials, and the bylines of anything you publish elsewhere. Consistency is boring, cheap, and among the highest-leverage hours in GEO.
Then make the basic facts machine readable. Organization schema on the homepage with a sameAs array pointing to every profile you control, a real About page that states in plain sentences what you do, for whom, where and since when, contact details that match everywhere they appear, and per-page structured data where it genuinely describes the content rather than decorating it. None of this forces a citation. What it does is remove ambiguity, and ambiguity is exactly what produces the confidently wrong descriptions of smaller businesses that AI answers have become known for.
Finally, exist in the reference sources these systems lean on. Retrieval favours pages that already rank and that are visibly maintained, which in practice means encyclopaedic references, established trade publications, and directories with checkable claims. Being listed and described accurately in a handful of those does more for entity clarity than a homepage redesign. If you sell services, a maintained profile in a directory of SEO agencies, carrying the same description as your own site and a link back to it, is one of the cheapest entity signals available to you.
- One canonical brand name, spelled and capitalised the same on every property
- Organization schema with sameAs links to every profile you control
- An About page stating what you do, for whom and where, in plain sentences
- The same short description reused across directories, profiles and press materials
- A single canonical domain, with redirects from the variants you own
Write paragraphs that can be quoted verbatim
Assistants quote paragraphs, not pages. The unit that survives retrieval is a self-contained block of forty to eighty words answering one question completely, with no dependency on the sentence before. If a paragraph opens with a pronoun pointing at an earlier section, or defers the answer to a table further down, it cannot be lifted without editing, and an engine with ten candidates will take one that needs none. Write each block as if it were the only thing anyone reads; the surrounding page adds depth for the humans who stay.
Answer before you elaborate. Put a direct, declarative response in the first sentence under each heading, then justify it. Define your terms explicitly, including the ones you assume everyone knows, because definitions are quoted far more often than arguments. Give every number its unit, its period and its source in the same sentence, since an unattributable figure is one a careful system avoids repeating. And keep headings question shaped wherever that is honest, because the closer a heading sits to the query the assistant formulated, the easier the match.
This is a writing discipline more than a technical one, and it changes what you should ask of whoever writes your pages. A content marketing partner working this way delivers text that reads slightly blunter than classic marketing copy: fewer narrative build-ups, more stated conclusions, more explicit scoping (this applies to this size of company, in this market, in 2026). The useful side effect is that these pages also perform better with impatient human readers, one of the rare cases where optimising for machines and for people agree.
- Answer in the first sentence, justify in the rest of the paragraph
- Keep each paragraph self-contained: no pronouns pointing at other sections
- Define your terms explicitly, even the obvious ones
- Attach a unit, a period and a source to every number you publish
- Use question-shaped headings when they match how buyers actually ask
- State the scope: who the advice applies to, and who it does not
Third-party validation beats your own website
For any query asking for a recommendation, a shortlist or a comparison, your own website is the weakest source in the room. Every vendor claims to be the best, and retrieval systems behave as though they know it: the pages fetched for best provider for my situation queries are overwhelmingly comparison articles, industry directories, review platforms and editorial roundups. You cannot write your way into that set from your own domain. You get there by being present, described accurately, and ideally verifiable, on the third-party pages that already rank for those questions.
Not every third-party page is worth the effort. The ones that matter share three traits: they rank for the queries your buyers actually ask, they are maintained rather than frozen three years ago, and their claims can be checked. That last trait is becoming the differentiator. A ranking built on verified performance data carries more weight than a list sold by placement, for humans and for the systems summarising them, so where such a ranking exists in your category, a place on the agency leaderboard beats ten guest posts.
Coverage should also be specific by geography and category, because that is how the questions arrive. People rarely ask for a good agency in the abstract: they ask for one in their country and their vertical, so presence on pages like SEO agencies in France or SEO agencies in the United States beats a generic worldwide listing. The same holds for software: a product listed in a maintained catalogue of SEO tools, with an accurate category and description, is far easier to place in a shortlist than one that exists only on its own site.
Why freshness matters more than in classic SEO
Freshness carries more weight here than in classic search, for a structural reason. When an assistant retrieves at answer time, it is trying to satisfy a question asked today, and it has visible signals about how old each candidate page is. In classic search, an authoritative page from 2021 can hold its position for years on accumulated reputation alone. Inside a retrieved set, that same 2021 page competing against an equivalent 2026 page tends to lose, particularly on anything where the world plausibly changed: prices, tools, regulations, market structure, best practice.
The practical consequence is that a small number of maintained pages beats a large archive of abandoned ones. Identify the pages that carry your commercial argument, put a visible and honest last-updated date on them, and revise them on a real cadence: quarterly for fast-moving subjects, twice a year for stable ones. Revising means changing the substance, adding what is new and deleting what is no longer true. Republishing identical text under a fresh date is transparent to readers and useless against systems that compare content rather than timestamps.
This also changes how you should think about archive depth. Pages that are no longer accurate are not neutral: they remain candidates for retrieval and will misrepresent you if fetched, while diluting the entity signal you spent months building. Auditing an archive once a year and deciding deliberately between update, consolidate and delete is unglamorous work that most content programmes skip. It is also one of the few GEO tasks with an immediate and observable effect, because you can watch which version of a page assistants start quoting within weeks.
How to measure AI visibility honestly
Measuring AI visibility honestly starts by admitting what does not exist. There is no Search Console for answer engines: no impression counts, no query volumes, no ranking positions, no complete log of the conversations where you were named. Referral data in your analytics captures only the fraction of conversations that end in a click, and attribution is inconsistent from one assistant to the next. Anyone presenting an AI visibility score to one decimal place is presenting a sample dressed up as a census. The honest alternative is sampling, done consistently over time.
Build a fixed set of thirty to sixty prompts that reflect how your buyers really ask: category questions, comparison questions, problem questions, and a few brand questions. Run them on a fixed schedule across the assistants your market actually uses, and record four things each time: whether you were mentioned, whether the description was accurate, which competitors appeared alongside you, and which sources were cited. That last field is the most actionable output of the entire exercise, because it names the third-party pages you should go and improve next.
Expect noise. The same prompt can produce different answers minutes apart, so read trends over months instead of reacting to single runs, and never compare one model against another as though they were the same instrument. Pair the sampling with what you can measure properly: assistant referral sessions isolated in your analytics setup, branded search volume, and direct traffic to the pages you worked on. None of these prove causation on their own, but together they tell you whether the programme is moving in the right direction.
- A fixed prompt set, rerun on a fixed schedule
- Mention, accuracy, competitor set and cited sources logged at every run
- Trends read over months, never over a single answer
- Assistant referral sessions isolated in your analytics
- Each model compared only against itself over time
What does not work
Three families of tactics are sold as GEO and do not work. The first is keyword stuffing rebranded for language models: repeating the brand name in every paragraph, appending lists of prompts you hope to appear for, hiding entity keywords in the footer. Retrieval runs on relevance and quality signals, and generation summarises meaning rather than counting terms, so the main measurable effect is a page that reads badly to the humans who arrive on it. The tactic fails on its own terms, before anyone asks whether it is legitimate.
The second is volume: publishing hundreds of generated pages to increase surface area. By construction, that text restates what models already contain, which gives a retrieval system no reason to prefer it and gives an assistant nothing quotable it could not have produced itself. It also carries real risk on the classic search side, which our piece on AI content and Google covers in detail. Generation is a legitimate drafting tool; it is not a substitute for having something specific, verifiable and yours to say.
The third is manipulation: instructions hidden in the page for the model, invisible text, content cloaked for AI crawlers. Beyond being deceptive, these are fragile, publicly documented and increasingly filtered, and they fail the simplest test of a durable strategy: could you describe the tactic to your own customers? If you are hiring for this work, ask what the candidate would still do if every model changed next quarter. The selection framework for SEO agencies applies here unchanged, and a verified agency directory is a reasonable place to start.