# AI crawlers and what an agent reads on your site

Fifteen documented AI crawlers do three different jobs, and one group of them decides whether an AI answer can cite you at all. This lesson reads our own run: what robots.txt allowed, what the agentic report could not measure, and which proposal the orchestrator threw out.

Claude SEO: do SEO inside Claude Code · Module 2 · 4 of 13 · 18 min

> This course is independent. It has no affiliation with Anthropic or with AgriciDaniel, who writes claude-seo, and nobody on their side reviewed it. Run with claude-seo v2.4.1 on 2026-10-03.

URL: https://seodraft.app/learn/claude-seo/ai-crawlers-and-agent-readiness
Learn: https://seodraft.app/learn/claude-seo.md

## What you will be able to do

- Name the three kinds of AI crawler and say what blocking each one costs you.
- Read your own robots.txt result per crawler, at the home page and at a content path.
- Tell a measurement the run could not take from an opportunity a draft standard proposes.

AI crawlers do three different jobs under one name, and the lines in your robots.txt decide which of the three you keep. One group decides whether an AI product can cite and link you. One opens a page because a person asked their assistant about it. One collects text for model training, and closing that door leaves the first two exactly where they were. This lesson reads our own run against this site on 2026-10-03: fifteen documented crawlers checked against the live robots.txt, what claude-seo's `agentic` and `geo` agents reported, the one measurement the run could not take, and the proposal the orchestrator threw out. The site is the one you are reading, so every figure below belongs to a report whose subject you can check yourself.

## The three kinds of AI bot, and what blocking each one costs

Blocking a crawler costs you something different depending on which of three jobs that crawler does. The check we ran groups the fifteen bots that way, and the group is the only thing that tells you what a `Disallow` buys.

| Group | What a block costs | The crawlers we checked |
| --- | --- | --- |
| Search | citations and links inside an AI product's answers | Googlebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, meta-webindexer |
| User request | the page an assistant opens because somebody asked | Claude-User, ChatGPT-User, Perplexity-User |
| Training | the text used to train a model | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot |

The effects are the operators' own words. Blocking Googlebot drops you out of Google Search and every Search feature, AI Overviews included. Blocking OAI-SearchBot keeps your pages out of ChatGPT search answers; Claude-SearchBot, out of Claude's search index; PerplexityBot, out of Perplexity's results and links; Applebot, out of Siri, Spotlight and Safari search; meta-webindexer, out of what Meta AI cites.

## Blocking a training crawler leaves search inclusion alone

Blocking GPTBot keeps your text out of OpenAI's training data and leaves ChatGPT search exactly where it was, because those are two tokens with two jobs. Every major operator splits them the same way:

- OpenAI: GPTBot trains, OAI-SearchBot answers in search.
- Anthropic: ClaudeBot trains, Claude-SearchBot indexes for Claude's search results.
- Google: Google-Extended covers Gemini training and grounding in the Gemini apps; Googlebot covers Search.
- Apple: Applebot-Extended covers foundation-model training; Applebot covers Siri, Spotlight and Safari.

CCBot is the odd one. It feeds Common Crawl, an open dataset many models train on, so blocking it reaches further than any single company. meta-externalagent is Meta's training crawler, and meta-webindexer is the token that decides whether Meta AI can cite you.

## What our robots.txt actually says

This site's robots.txt has a single `*` group, and all fifteen crawlers inherit its `Allow`. The file answered 200, declared a sitemap, and carried no `Content-Signal` directives on any group. seodraft's check read it at the home page and at `/blog/` on 2026-10-03 and came back with fifteen allowed bots and an empty findings list.

Two of those allowances mean less than they look. ChatGPT-User and Perplexity-User are recorded with `honorsRobots: false`, because OpenAI documents that robots.txt rules may not apply to a fetch a user asked for, and Perplexity documents that Perplexity-User generally ignores robots.txt. Allowing them changes nothing, and a `Disallow` aimed at them is a request their operators say they can decline. Enforcement at that level is a firewall's job.

## llms.txt is published here, and Google says Search does nothing with it

Our `/llms.txt` answers 200 with 6,896 bytes of plain text: an H1, a summary and six sections pointing at the pages worth quoting. Google's own AI optimization guide added a mythbusting note on 2026-06-15 saying llms.txt is not needed for Google Search and that it neither helps nor hurts a site there ([developers.google.com](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide)).

Both facts hold at once. We keep the file because it costs nothing to maintain and it tells an agent reading by hand where the dated, citable page lives. Nobody should expect a ranking from it, and claude-seo's `agentic` agent files it as a P3 item for the same reason.

## The agentic report named one measurement it could not take

Lighthouse Agentic Browsing came back unmeasured in our run: PageSpeed Insights answered 429, quota exceeded, on both the mobile and the desktop strategy, because the session had no API key. The report printed "fraction unknown — cannot report X/N" where the number goes.

That blank is the useful part. A section with no measurement is a hole in your evidence, and the scores around it were never built on it, so reading the hole first stops you from quoting a pass that was never issued. The fix is an API key and a rerun; until then, write "unmeasured" in your notes and leave the row empty.

## The 100/100 is a heuristic over the accessibility tree

The agent-UX score of 100/100 in our run came from counting the rendered page's accessibility tree, so it reports exactly what it counted. Ours: 503 nodes, 71 interactive, 0 unnamed interactive elements, 0 inputs missing a label or an `aria` name, 87 real `<a>`, 3 real `<button>`, 0 `div`-with-onclick widgets and 16 semantic landmarks. It flagged no partial-render or JavaScript-shell problem, which is what server rendering is supposed to produce.

What the score tells you is that an agent reading your HTML finds named controls and real landmarks instead of a pile of unlabelled divs. What it leaves open is whether an agent can finish a task on your site. Treat it as a floor you have cleared.

## Draft proposals are opportunities, not defects

Four items our `agentic` agent raised are proposals no standard has merged, and the report files them with the date each status was checked:

- `/.well-known/ai-catalog.json` returns 404. Draft, checked 2026-09-23.
- `/.well-known/agent-card.json` returns 404. The MCP Server Card proposal, SEP-2127, was unmerged as of 2026-09-12.
- Markdown delivery is absent: no `Vary: Accept`, no `X-Markdown-Tokens`, `/index.md` 404. No primary source shows a consumer agent sending `Accept: text/markdown`.
- WebMCP markup is absent. It is a Chrome origin trial, M149 to M156, with no ship date.

`Content-Signal` sits in the same bucket. Our robots.txt declares none, and the report calls that informational, since Content-Signal is a Cloudflare CC0 policy whose IETF draft expired on 2026-04-04. The two `.well-known` files that did answer, `oauth-protected-resource` and `oauth-authorization-server`, returned 200 with valid JSON.

An item dated like that is a decision you can make on your own calendar. When the status changes, the date in the report tells you how old the advice has become.

## When two agents disagree, the orchestrator decides

Our `geo` agent scored 82/100 and proposed adding FAQPage markup to `/ai-info` and the feature pages; the orchestrator discarded the proposal. Google retired FAQ rich results on 2026-05-07, so the markup wins no SERP feature, and an afternoon spent adding it would have bought nothing. The FAQPage already on those pages stays where it is, as information for anything that reads the JSON-LD.

That is the same habit [lesson 3](/learn/claude-seo/seo-audit-with-claude) asked you to build: read the verification section before you open a file. A sub-agent reports what it saw in its own slice, and the orchestrator is the pass that compares the slice with the live site. Your report will carry both halves, and the disagreement is where the reading is.

## What to change in your robots.txt and what to leave alone

Change a line only where the block costs you something you want back. Four decisions cover almost any site:

- Keep the search group allowed if you want AI answers to cite and link you. A block there is the only one that removes you from an answer.
- Decide training on your own licensing view. Blocking GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent or CCBot affects training data, and the search tokens above keep working.
- Skip the `Disallow` aimed at ChatGPT-User or Perplexity-User. Their operators document that it can be ignored, so if you need it enforced, enforce it at the edge.
- Check your CDN separately. Cloudflare has blocked AI crawlers by default on new domains since 2025-07-01, and robots.txt cannot show you a firewall rule. The check raises a warning when it sees a Cloudflare edge for exactly that reason.

Rerun the check after any edit, and rerun it when you move hosts. With Claude Code closed, [the free AI crawler check](/tools/ai-crawlers) reads the same registry under the same RFC 9309 rules, with no install and no account.

## Steps

### Step 1 · Read your site the way an agent reads it (claude-seo)

claude-seo's README documents this as a standalone command. Ours ran inside `/seo audit`, which starts the `agentic` agent among eleven, so what we quote is that agent's own file rather than a separate session. Either way the agent fetches your pages, renders one, counts its accessibility tree and probes the `.well-known` paths an agent would look for.

```
/seo agentic https://your-site.com
```

What you should see: A Lighthouse Agentic Browsing section, an agent-UX heuristic over the accessibility tree, and a list of probes with their status codes. Ours scored 100/100 on the heuristic and left the Lighthouse section unmeasured, because PageSpeed answered 429 with no API key configured.

### Step 2 · Score how quotable your pages are (claude-seo)

Same note: the README documents the standalone command and ours ran as part of `/seo audit`. The `geo` agent asks whether an answer engine can quote you — how self-contained your answers are, how the headings are built, which entity signals exist, and whether the text is in the HTML before JavaScript runs.

```
/seo geo https://your-site.com
```

What you should see: A readiness score with its dimensions spelled out. Ours was 82/100: citability 20/25, structural readability 16/20, multi-modal content 9/15, authority and brand signals 17/20, technical accessibility 20/20.

### Step 3 · Get the robots.txt answer per crawler (seodraft)

`check_ai_access` is free and read-only over MCP. It reads your robots.txt under RFC 9309 rules and reports every crawler whose operator documents the token and what blocking it does, grouped by what the block costs you. There is no score, and a crawler that says it ignores robots.txt is reported as saying so.

```
Run check_ai_access on my site and show me every crawler at the home page and at a content path.
```

What you should see: One row per crawler for both paths. Ours came back on 2026-10-03 with 15 bots, all allowed by the `*` group, robots.txt at 200 with a declared sitemap, llms.txt present and an empty findings list.

### Step 4 · Separate the drafts from the measurements (claude-seo)

Paste this as a prompt in the same session. An agentic report mixes three kinds of line: what the run measured, what it failed to measure, and what some standards body has proposed and nobody has merged. Only the first kind is a result about your site.

```
From the agentic report, list every item whose status is a draft or a proposal with the date it was checked, and every measurement the run could not take.
```

What you should see: Two short lists. Ours held four draft-status opportunities — `ai-catalog.json`, `agent-card.json`, Markdown delivery and WebMCP — and one measurement the run could not take, Lighthouse Agentic Browsing.

## Checklist

- [ ] I can name the three crawler groups and what blocking each one costs.
- [ ] I have my robots.txt answer per crawler, at the home page and at a content path.
- [ ] I know which of my blocks touch citations and which touch training alone.
- [ ] I wrote down what my agentic run could not measure, and why.
- [ ] I decided which robots.txt lines to change and which to leave alone.

## Checkpoint

### You block GPTBot in robots.txt. Does ChatGPT stop citing your pages?

No. GPTBot collects training data, and ChatGPT search reads through OAI-SearchBot, which has its own line in robots.txt and its own answer. Every big operator splits the two tokens the same way: ClaudeBot trains while Claude-SearchBot indexes, Google-Extended feeds Gemini while Googlebot feeds Search, Applebot-Extended trains while Applebot serves Siri and Spotlight. Blocking training is a licensing decision, and search inclusion is decided by the search tokens.

### Your agentic report has an empty Lighthouse Agentic Browsing section. What do you write down?

That it was unmeasured, and the reason. Ours said PageSpeed Insights answered 429, quota exceeded, on both the mobile and the desktop strategy with no API key configured, so the report printed "fraction unknown — cannot report X/N" where the number goes. An empty measurement is a hole in your evidence; the scores printed beside it were built from other checks. Configure a key, rerun, and read that section before you treat it as a pass.

## What this lesson said

- AI crawlers do three jobs: citations in an AI product's answers, fetches a person asked for, and model training. The group is what tells you the cost of a block.
- Blocking a training crawler leaves search inclusion alone. GPTBot and OAI-SearchBot, ClaudeBot and Claude-SearchBot, Google-Extended and Googlebot are pairs with separate jobs.
- Our robots.txt has one `*` group and all 15 documented crawlers inherit its `Allow`. ChatGPT-User and Perplexity-User report that they may ignore robots.txt, so those two allowances change nothing.
- Our `agentic` run scored 100/100 on the accessibility-tree heuristic over 503 nodes and 71 interactive elements, and left Lighthouse Agentic Browsing unmeasured after PageSpeed answered 429 with no API key.
- The `geo` agent scored 82/100 and proposed FAQPage markup; the orchestrator discarded it, because Google retired FAQ rich results on 2026-05-07.

## Questions

### Should I block the AI training crawlers?

That is your call about licensing, and the check never frames it as a problem. Blocking GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent or CCBot keeps your text out of those training sets and leaves every search token working as before. The one thing worth knowing first is that CCBot feeds Common Crawl, an open dataset many models train on, so that block reaches further than any single company.

### Do I need an llms.txt?

Google says you do not need one for Google Search, and that it neither helps nor hurts there; the mythbusting note went into its AI optimization guide on 2026-06-15. This site publishes one anyway, 6,896 bytes with a summary and six sections, because it points an agent reading by hand at the dated page we would rather be quoted from. Publish it if that reason applies to you, and expect nothing from it in Search.

### Can I check this without installing anything?

Yes. The free AI crawler check at /tools/ai-crawlers runs the same engine as seodraft's `check_ai_access`: the same registry of documented bots, the same RFC 9309 reading of robots.txt, the home page and a content path, no sign-in and no score. It is the fastest way to look at a site you do not own, or to show somebody what their robots.txt is doing.

Next lesson: [Keyword research with Claude](https://seodraft.app/learn/claude-seo/keyword-research-with-claude)
