Updated nightly. Method published in full. Dataset open.
The open index of how the web treats AI agents.
We measure the most-visited sites on the web and publish what we find. Which AI crawlers each site blocks. Whether it publishes llms.txt or agents.md. Whether it quietly serves crawlers something different from what it serves you. And which platforms and CDNs are making that decision on the operator's behalf.
What the last crawl found
3,678 domains measured on 2026-08-09. Sites we could not reach are excluded rather than counted as failures.
Who is actually deciding
Most operators never formed a policy on AI crawlers. Their edge network did, by default, and they inherited it. Sites behind fastly block an answer-surface crawler 24.8% of the time, against 3.2% behind google.
| Edge network | Sites | Blocking AI | Proportion blocking | Mean score |
|---|---|---|---|---|
| Cloudflare | 1,076 | 188 (17.5%) | 71.7 | |
| Amazon CloudFront | 424 | 87 (20.5%) | 63.6 | |
| Akamai | 307 | 44 (14.3%) | 63.8 | |
| Fastly | 222 | 55 (24.8%) | 65.9 | |
| Google Cloud / GFE | 154 | 5 (3.2%) | 61 | |
| Vercel | 48 | 3 (6.3%) | 76.5 |
Readiness by platform
What a site is built on predicts how legible it is to an agent.
| Platform | Sites | Blocking AI | Proportion blocking | Mean score |
|---|---|---|---|---|
| WordPress | 298 | 52 (17.4%) | 72.3 | |
| Next.js | 295 | 48 (16.3%) | 67.3 | |
| Adobe Experience Manager | 156 | 6 (3.8%) | 67.6 | |
| Drupal | 120 | 4 (3.3%) | 64.3 | |
| HubSpot CMS | 78 | 3 (3.8%) | 74.2 | |
| Nuxt | 57 | 4 (7.0%) | 63.6 |
Which crawlers get shut out
Share of measured sites whose robots.txt denies each crawler the site root.
| Crawler | Operator | Blocked by | Proportion |
|---|---|---|---|
| CCBot | Common Crawl | 585 (15.9%) | |
| GPTBot | OpenAI | 549 (14.9%) | |
| Bytespider | ByteDance | 548 (14.9%) | |
| ClaudeBot | Anthropic | 529 (14.4%) | |
| meta-externalagent | Meta | 493 (13.4%) | |
| Google-Extended | 478 (13.0%) | ||
| Applebot-Extended | Apple | 456 (12.4%) | |
| Amazonbot | Amazon | 445 (12.1%) | |
| Diffbot | Diffbot | 387 (10.5%) | |
| cohere-ai | Cohere | 386 (10.5%) | |
| PerplexityBot | Perplexity | 383 (10.4%) | |
| YouBot | You.com | 331 (9.0%) |
Least agent-ready right now
Fully measured sites with the lowest scores. Partial assessments are excluded because a renormalised score is not comparable to a complete one.
| Rank | Domain | Score | Answer-surface crawlers | Agent files |
|---|---|---|---|---|
| 1228 | furaffinity.netrefused GPTBot | Score 11 out of 100, grade F | 11 blocked | none |
| 56 | tiktok.com | Score 15 out of 100, grade F | 11 blocked | none |
| 501 | amazon.co.jp | Score 15 out of 100, grade F | 10 blocked | none |
| 612 | amazon.fr | Score 15 out of 100, grade F | 10 blocked | none |
| 734 | amazon.es | Score 15 out of 100, grade F | 10 blocked | none |
| 3508 | amazon.ae | Score 15 out of 100, grade F | 10 blocked | none |
| 25 | amazon.com | Score 18 out of 100, grade F | 10 blocked | none |
| 317 | amazon.co.uk | Score 18 out of 100, grade F | 10 blocked | none |
| 382 | amazon.de | Score 18 out of 100, grade F | 10 blocked | none |
| 642 | amazon.ca | Score 18 out of 100, grade F | 10 blocked | none |
Why this exists
Publishers are deciding, one robots.txt at a time, whether AI systems may read the web. Those decisions are made quietly, changed without announcement, and are individually trivial to check but collectively invisible.
CrawlIndex checks them on a schedule and keeps the receipts. The rubric is published, every score is arithmetic over archived evidence, and no language model touches the numbers. The whole dataset is downloadable. If you disagree with a result you can read exactly how it was reached and recompute it yourself.
Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "The state of AI crawler access." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.