Ask how to get cited by ChatGPT and most advice starts in the same place: open robots.txt, check you're not blocking GPTBot or OAI-SearchBot, move on. It costs nothing to check, so it's not bad advice. It's just very rarely the actual problem. We read 197 real production websites for our own AI SEO audit, run between 31 July and 6 September 2026, and not one of them blocked a crawler that feeds an AI answer in robots.txt. If your brand isn't showing up in ChatGPT, Perplexity or Google's AI answers, the block probably isn't there, and it's worth knowing what to check in its place.
Is robots.txt blocking you from being cited?
Almost certainly not. We measured this directly: across the 197 sites in our own audit data, zero blocked any of the four crawlers that decide whether a page can be quoted in an AI answer at all: Googlebot (AI Overviews and AI Mode), OAI-SearchBot (ChatGPT), Claude-SearchBot and PerplexityBot. Six sites blocked a training crawler such as GPTBot or CCBot, a separate and entirely legitimate choice about training data rather than citation. Our access check scores 100 out of 100 on 67.5% of sites, and only 11 of the 197, 5.6%, raised any fatal access finding at all. For almost everyone, robots.txt was never where the problem lived.
The block that robots.txt can't show you
The problem, when there is one, tends to sit somewhere a robots.txt read can't see. We ran our own probe of 70 of those sites, requesting the home page as each of the four named crawlers directly, rather than just reading what the file said was allowed. Eight of the 70, 11.4%, refused at least one named AI crawler at the edge, meaning a firewall or CDN rule turned the request away before robots.txt was ever consulted. Six of those eight refused all four crawlers alike, answer crawlers and the training crawler together, which points to one bot-management switch rather than four separate decisions about ChatGPT, Perplexity, Claude and Google. None of the eight said so anywhere a visitor could read. A tool that only checks robots.txt, which is most of them, would report all eight as wide open.
Two things worth knowing about our own test before you read too much into it. We send the user agent string each provider publishes, so the probe identifies itself honestly as a named crawler; a bot-management product can also verify by IP range, so a refusal to us is evidence about the rule rather than proof the operator's own crawler is refused the same way. And the 70 sites we probed skew towards travel operators and small businesses, so treat that 11.4% as a panel reading, not a population figure.
The silence is the part worth sitting with. A site that deliberately blocks GPTBot in robots.txt is making a call, visible to anyone who reads the file, and it's easy to write "we don't allow AI training" in a page footer if that's the intent. A firewall rule catching an answer crawler by accident produces no such trace: the operator doesn't know, the crawler doesn't retry, and every free visibility checker built to read robots.txt comes back clean. The only way anyone finds out is by testing what actually happens to a request from that specific crawler, which is precisely the step a robots.txt read skips.
How do you check your own site for this?
Test what a named crawler receives, not just what your robots.txt permits. Send a request with the crawler's own published user agent and look at the response code and body, the same way our probe did. A firewall block returns a different answer from a robots.txt disallow, and only one of the two shows up if you just read the file.
Check your CDN or firewall's bot-management settings directly, separately from robots.txt. Cloudflare, for one, ships a named "AI Scrapers and Crawlers" rule that a lot of teams switch on to keep AI training crawlers off their content, then forget it also catches every other agent with "AI" or "bot" in its name, answer crawlers included. If a rule refuses one named crawler, check whether it refuses the rest too: in our probe, most of the sites that blocked anything blocked all four identically, one setting rather than four intentional calls.
Whoever owns the CDN account is usually a different person from whoever owns the content, which is part of why this goes unnoticed. A developer or an IT team switches on a bot-management default to stop scraping, marketing never hears about it, and the two teams have no shared dashboard where "we're invisible to ChatGPT" and "we changed a firewall rule six months ago" would ever sit next to each other. Asking the question directly, in a channel where both teams can answer, is sometimes the fastest fix of all: not a technical change, just a conversation nobody had thought to have.
Or skip the manual work. Our free AI SEO audit requests your pages as each of those four crawlers and reports what came back, not just what your robots.txt says should be allowed. It's the one check a robots.txt read on its own can't do for you.
How to get cited by ChatGPT once access isn't the problem

The report above is the shape of what an audit like this returns: one score out of 100, then four dials underneath it, worst finding first. Getting past the access dial turns out to be the easy part for nearly everyone. The same audit measured what a crawler finds once it's inside, across four dials: access, chunking (does a section still make sense read alone), length (is a section substantial enough to be worth indexing) and evidence (does it carry a number, a named source or a link).
Access is the one that's fine on almost every site. Evidence is the one that fails worst: 84.6% of sites scored under 40 out of 100 on it, meaning most sections have nothing checkable in them at all, as we found running the same check across a separate 164-site sample. Length is close behind, failing on 54.9% of sites. A crawler that reaches a paragraph with no number, no named source and no link has nothing to quote, however clean your access setup is.
That's the same ground the fuller list of what earns a citation covers in more depth: put a specific figure, a named source or an outbound link in every section that matters, keep sections long enough to stand alone, and write the answer to the obvious question in the first sentence rather than building up to it.
What to check this week
In order, because each step only matters once the one before it is ruled out:
- Confirm
robots.txtisn't blockingGPTBot,OAI-SearchBot,ClaudeBot,Claude-SearchBot,PerplexityBotorGoogle-Extended. Quick, and rarely where the fault is. - Send a request as one of those named crawlers, or run the free audit, and check what comes back at the firewall or CDN level, separately from robots.txt.
- If that's clean too, look at what the crawler finds once it's inside: does your most important page state a specific number, name a real source, or link out to one, in every section that matters?
Most sites we've read pass the first two checks without changing anything. The third is where nearly everyone has work to do, and it's also the one entirely within your own control: nobody else's crawler policy, nobody else's firewall, just what you put on the page. What AI search visibility actually measures is the place to start if appearance rate, share of AI voice and citation are new terms, and Babel42's free plan tracks one AI Buyer across any two of seven assistants every week if you'd rather watch the number move than check it once by hand.


