Answer engine optimisation

What actually gets you cited, in order

Almost every article about AEO is a list in its author's preferred order. This one is not ours. It follows a published meta-analysis of 54 experiments, patents and case studies that scored each factor on three things: how often the finding repeated, how strong the evidence was, and whether an operator's own documentation or a patent backs it.

We have kept the three findings that are inconvenient for anybody selling AEO work, including us: structured data scores 5.6 out of 10, llms.txt scores 2.0, and adding citations measurably hurts pages that already rank first.

Four questions, and they are asked in order

An assistant has to be able to fetch your page, decide it answers the question, find a section worth quoting, and have some reason to trust you. A failure at any step makes everything after it irrelevant, which is why a list sorted by difficulty or by novelty sends people to the wrong job first.

  1. Can it be read

    4 factors · best evidenced: url accessibility

    Whether an assistant can fetch the page and use what it finds. Nothing else on this list matters until these pass.

  2. Does it match the question

    5 factors · best evidenced: search rank

    Whether this page is the one that answers what was asked, including the related questions asked alongside it.

  3. Is it written to be quoted

    8 factors · best evidenced: answer near the top

    Whether a single section, handed over on its own, answers something. This is the half most sites have never been edited for.

  4. Is the source trusted

    6 factors · best evidenced: topic cluster ranking

    What the model already believes about you. Mostly earned elsewhere, and mostly not visible on your own site.

All 23, worst evidence last

The score is how well supported the factor is, not how hard it is to fix. Our audit measures 16 of these on your own site. The rest are off it: what your page ranks for, and what the model already believes about you.

  1. 9.5

    #1

    URL accessibility

    our audit checks this

    Whether the page can be fetched at all when an assistant goes looking.

    Let the answer-feeding crawlers in, at robots.txt and at your firewall, and make sure the page's content is in the HTML your server sends. This is first for a reason: nothing below it can happen if this fails.

  2. 9.4

    #2

    Search rank

    not on your site

    How the page ranks in ordinary search for the question being asked.

    Conventional SEO, essentially unchanged. Assistants lean heavily on existing search results, which is the least fashionable finding in this whole list and one of the best supported.

  3. 9.3

    #3

    Fan-out rank

    not on your site

    How the page ranks for the related questions an assistant asks alongside the typed one.

    Cover the sub-questions around your topic, not only the headline one. Google states its AI features issue multiple related searches at once, so you compete for questions nobody typed.

  4. 9.2

    #4

    Preview controls

    our audit checks this

    Whether nosnippet, max-snippet or data-nosnippet stop your text being shown.

    Remove them from pages you want quoted. This one is quietly common: a page can be perfectly indexed and still forbidden to appear.

  5. 9.2

    #5

    Query-answer match

    our audit checks this

    Whether the page actually answers the question, in the words it was asked in.

    Write the answer plainly, near the top, in the language of the question rather than of your product.

  6. 9.0

    #6

    Intent-format match

    not on your site

    Whether the page type suits the question. A comparison question wants a comparison page.

    Match the shape to the ask: a list for "best", a walkthrough for "how", a definition for "what is".

  7. 8.9

    #7

    Topic cluster ranking

    not on your site

    Whether the site ranks across a group of related questions rather than one.

    Build depth on a topic instead of one page per keyword. Off-site to measure, so this audit cannot see it.

  8. 8.8

    #8

    Answer near the top

    our audit checks this

    How far down the page the substance starts.

    Answer in the first paragraph, then explain. Two sentences under the heading is enough.

  9. 8.6

    #9

    AI-ready structure

    our audit checks this

    Headings, lists and tables a machine can lift a specific fact out of.

    Use real headings, real lists and real tables. A comparison drawn as a picture cannot be quoted.

  10. 8.3

    #10

    Factually specific

    our audit checks this

    Whether the page contains checkable facts rather than description.

    Put the number in. Assistants quote specifics and skip adjectives.

  11. 8.1

    #11

    Explicit phrasing

    our audit checks this

    Whether claims are stated plainly rather than implied.

    Say "costs $9.99 a month" rather than "affordable pricing". The vague version cannot be quoted as an answer.

  12. 8.0

    #12

    Cites sources

    our audit checks this

    Whether facts are backed by a reference.

    Link where a claim came from. ⚠️ The published experiment behind this found it helps pages ranking poorly and slightly hurts pages already ranked first, so spend the effort where you are weak.

  13. 8.0

    #13

    Self-contained passages

    our audit checks this

    Whether a section still makes sense handed over on its own.

    Open each section by naming its subject rather than with "this", "it" or "however". The section arrives without the one above it.

  14. 7.6

    #14

    Content visibility

    our audit checks this

    Whether the important text is in the HTML rather than in an image or behind JavaScript.

    Render on the server, and write in text what a picture is currently saying.

  15. 7.0

    #15

    Freshness

    our audit checks this

    How current the page is, and whether it says so.

    Show a real updated date and keep your sitemap dates honest. Stamping today on everything teaches crawlers to ignore the file.

  16. 6.8

    #16

    Brand and entity trust

    not on your site

    What the model already knows about you before it reads anything.

    Earned coverage, mentions and reviews elsewhere. This is off your site entirely, which is why no site audit can measure it.

  17. 6.7

    #17

    Length

    our audit checks this

    Whether sections are long enough to answer and short enough to stay clear.

    Aim for sections of roughly 100 to 300 words. ⚠️ That band is prevailing practice in retrieval systems, not a measured optimum.

  18. 6.3

    #18

    Language

    our audit checks this

    Whether the page is in the language of the question.

    Declare it: <html lang="en-GB">. Cheap, and it is skipped constantly.

  19. 5.8

    #19

    Entity consistency

    our audit checks this

    Whether you name your products, brand and people the same way everywhere.

    Pick one name per thing and use it, including in the sections you most want quoted.

  20. 5.6

    #20

    Structured data

    our audit checks this

    Schema markup identifying what the page and the organisation are.

    ⚠️ Worth knowing before you plan a sprint: Google states there is no special structured data needed for its AI features, and this scores near the bottom of the list. Add Organization and Article schema, then move on.

  21. 5.4

    #21

    Known source

    not on your site

    Whether the URL was already known to the model from training.

    Nothing directly actionable, and nothing an audit can see.

  22. 5.0

    #22

    Domain authority

    not on your site

    Link-based popularity, as the SEO industry measures it.

    The usual long game. Note how far down this sits: it matters less here than it does in search.

  23. 2.0

    #23

    llms.txt

    our audit checks this

    A machine-readable file some people publish for AI tools.

    Skip it. Across 137,000 domains, 28% publish one and 97% of those files received no requests at all in a month, and Google says its Search does not use them. We report whether you have one and never count its absence against you.

Scores from Cyrus Shepard's AI citation ranking factors analysis, published 7 May 2026, which synthesises 54 experiments, patents and case studies. We read the figures from that analysis directly. The commentary under each factor is ours.

Start with the ones you can see

Everything in the first group, and most of the third, is measurable on your own site right now. The free audit fetches your pages the way an AI crawler does and reports what came back, in a few seconds and without an account.

Questions people actually ask

What is AEO?

AEO, or answer engine optimisation, is the work of making a website usable by AI assistants rather than only by search engines. The difference that matters: a search engine sends someone a link to your page, while an assistant reads a section of your page and answers in its own words. That makes the section, not the page, the unit that competes, and it makes being readable a prerequisite rather than a detail.

Is AEO different from SEO?

Less than most articles suggest. The second best-evidenced factor in this list is ordinary search rank, because assistants lean heavily on existing search results. AEO adds two things on top: whether a crawler can read your pages at all, since most AI crawlers never run JavaScript, and whether each section stands up alone when it is handed over without the rest of the page.

Do I need an llms.txt file?

No, on the current evidence. Across 137,000 domains with traffic, 28% publish an llms.txt file and 97% of those files received no requests at all in a month. Google states its Search does not use them. It scores 2.0 out of 10 here, the lowest of the 23 factors, and Babel42's audit reports whether you have one without ever counting its absence against you.

Does structured data help with AI search?

A little, and less than the amount of advice about it suggests. Google states plainly that there is no special schema.org structured data needed to appear in its AI features, and the evidence-weighted score here is 5.6 out of 10, near the bottom. Add Organization and Article schema because it helps elsewhere, then spend the remaining time on whether your pages answer the question near the top.

Does blocking GPTBot remove me from ChatGPT?

No, and this is the most common mistake in AI SEO. GPTBot collects training data. OAI-SearchBot is the crawler that feeds ChatGPT's search answers, and OpenAI states that sites opted out of it will not be shown there. Blocking training crawlers while allowing answer crawlers is a coherent position that thousands of publishers hold deliberately.

Does page speed affect AI citations?

There is no good evidence that it does. Google publishes no speed requirement for AI Overviews or AI Mode, and the largest analysis of it, covering roughly 107,000 pages appearing in AI Overviews, found correlations close to zero. Speed matters for your visitors and for conventional search. Treat any AEO checklist that leads with Core Web Vitals with suspicion.

How do I know if AI assistants can read my site?

Run an audit that fetches your pages the way a crawler does rather than the way a browser does. Babel42's free AI SEO audit reads your robots.txt against the crawlers that feed AI answers, requests your pages as an identified crawler, checks what survives without JavaScript, and measures whether each section can be understood on its own.