AI crawlers do three different jobs under one name, and the lines in your robots.txt decide which of the three you keep. One group decides whether an AI product can cite and link you. One opens a page because a person asked their assistant about it. One collects text for model training, and closing that door leaves the first two exactly where they were. This lesson reads our own run against this site on 2026-10-03: fifteen documented crawlers checked against the live robots.txt, what claude-seo's agentic and geo agents reported, the one measurement the run could not take, and the proposal the orchestrator threw out. The site is the one you are reading, so every figure below belongs to a report whose subject you can check yourself.
The three kinds of AI bot, and what blocking each one costs
Blocking a crawler costs you something different depending on which of three jobs that crawler does. The check we ran groups the fifteen bots that way, and the group is the only thing that tells you what a Disallow buys.
| Group | What a block costs | The crawlers we checked |
|---|---|---|
| Search | citations and links inside an AI product's answers | Googlebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, meta-webindexer |
| User request | the page an assistant opens because somebody asked | Claude-User, ChatGPT-User, Perplexity-User |
| Training | the text used to train a model | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot |
The effects are the operators' own words. Blocking Googlebot drops you out of Google Search and every Search feature, AI Overviews included. Blocking OAI-SearchBot keeps your pages out of ChatGPT search answers; Claude-SearchBot, out of Claude's search index; PerplexityBot, out of Perplexity's results and links; Applebot, out of Siri, Spotlight and Safari search; meta-webindexer, out of what Meta AI cites.
Blocking a training crawler leaves search inclusion alone
Blocking GPTBot keeps your text out of OpenAI's training data and leaves ChatGPT search exactly where it was, because those are two tokens with two jobs. Every major operator splits them the same way:
- OpenAI: GPTBot trains, OAI-SearchBot answers in search.
- Anthropic: ClaudeBot trains, Claude-SearchBot indexes for Claude's search results.
- Google: Google-Extended covers Gemini training and grounding in the Gemini apps; Googlebot covers Search.
- Apple: Applebot-Extended covers foundation-model training; Applebot covers Siri, Spotlight and Safari.
CCBot is the odd one. It feeds Common Crawl, an open dataset many models train on, so blocking it reaches further than any single company. meta-externalagent is Meta's training crawler, and meta-webindexer is the token that decides whether Meta AI can cite you.
What our robots.txt actually says
This site's robots.txt has a single * group, and all fifteen crawlers inherit its Allow. The file answered 200, declared a sitemap, and carried no Content-Signal directives on any group. seodraft's check read it at the home page and at /blog/ on 2026-10-03 and came back with fifteen allowed bots and an empty findings list.
Two of those allowances mean less than they look. ChatGPT-User and Perplexity-User are recorded with honorsRobots: false, because OpenAI documents that robots.txt rules may not apply to a fetch a user asked for, and Perplexity documents that Perplexity-User generally ignores robots.txt. Allowing them changes nothing, and a Disallow aimed at them is a request their operators say they can decline. Enforcement at that level is a firewall's job.
llms.txt is published here, and Google says Search does nothing with it
Our /llms.txt answers 200 with 6,896 bytes of plain text: an H1, a summary and six sections pointing at the pages worth quoting. Google's own AI optimization guide added a mythbusting note on 2026-06-15 saying llms.txt is not needed for Google Search and that it neither helps nor hurts a site there (developers.google.com).
Both facts hold at once. We keep the file because it costs nothing to maintain and it tells an agent reading by hand where the dated, citable page lives. Nobody should expect a ranking from it, and claude-seo's agentic agent files it as a P3 item for the same reason.
The agentic report named one measurement it could not take
Lighthouse Agentic Browsing came back unmeasured in our run: PageSpeed Insights answered 429, quota exceeded, on both the mobile and the desktop strategy, because the session had no API key. The report printed "fraction unknown — cannot report X/N" where the number goes.
That blank is the useful part. A section with no measurement is a hole in your evidence, and the scores around it were never built on it, so reading the hole first stops you from quoting a pass that was never issued. The fix is an API key and a rerun; until then, write "unmeasured" in your notes and leave the row empty.
The 100/100 is a heuristic over the accessibility tree
The agent-UX score of 100/100 in our run came from counting the rendered page's accessibility tree, so it reports exactly what it counted. Ours: 503 nodes, 71 interactive, 0 unnamed interactive elements, 0 inputs missing a label or an aria name, 87 real <a>, 3 real <button>, 0 div-with-onclick widgets and 16 semantic landmarks. It flagged no partial-render or JavaScript-shell problem, which is what server rendering is supposed to produce.
What the score tells you is that an agent reading your HTML finds named controls and real landmarks instead of a pile of unlabelled divs. What it leaves open is whether an agent can finish a task on your site. Treat it as a floor you have cleared.
Draft proposals are opportunities, not defects
Four items our agentic agent raised are proposals no standard has merged, and the report files them with the date each status was checked:
/.well-known/ai-catalog.jsonreturns 404. Draft, checked 2026-09-23./.well-known/agent-card.jsonreturns 404. The MCP Server Card proposal, SEP-2127, was unmerged as of 2026-09-12.- Markdown delivery is absent: no
Vary: Accept, noX-Markdown-Tokens,/index.md404. No primary source shows a consumer agent sendingAccept: text/markdown. - WebMCP markup is absent. It is a Chrome origin trial, M149 to M156, with no ship date.
Content-Signal sits in the same bucket. Our robots.txt declares none, and the report calls that informational, since Content-Signal is a Cloudflare CC0 policy whose IETF draft expired on 2026-04-04. The two .well-known files that did answer, oauth-protected-resource and oauth-authorization-server, returned 200 with valid JSON.
An item dated like that is a decision you can make on your own calendar. When the status changes, the date in the report tells you how old the advice has become.
When two agents disagree, the orchestrator decides
Our geo agent scored 82/100 and proposed adding FAQPage markup to /ai-info and the feature pages; the orchestrator discarded the proposal. Google retired FAQ rich results on 2026-05-07, so the markup wins no SERP feature, and an afternoon spent adding it would have bought nothing. The FAQPage already on those pages stays where it is, as information for anything that reads the JSON-LD.
That is the same habit lesson 3 asked you to build: read the verification section before you open a file. A sub-agent reports what it saw in its own slice, and the orchestrator is the pass that compares the slice with the live site. Your report will carry both halves, and the disagreement is where the reading is.
What to change in your robots.txt and what to leave alone
Change a line only where the block costs you something you want back. Four decisions cover almost any site:
- Keep the search group allowed if you want AI answers to cite and link you. A block there is the only one that removes you from an answer.
- Decide training on your own licensing view. Blocking GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent or CCBot affects training data, and the search tokens above keep working.
- Skip the
Disallowaimed at ChatGPT-User or Perplexity-User. Their operators document that it can be ignored, so if you need it enforced, enforce it at the edge. - Check your CDN separately. Cloudflare has blocked AI crawlers by default on new domains since 2025-07-01, and robots.txt cannot show you a firewall rule. The check raises a warning when it sees a Cloudflare edge for exactly that reason.
Rerun the check after any edit, and rerun it when you move hosts. With Claude Code closed, the free AI crawler check reads the same registry under the same RFC 9309 rules, with no install and no account.