crawlindex

Updated nightly. Method published in full. Dataset open.

The open index of how the web treats AI agents.

We measure the most-visited sites on the web and publish what we find. Which AI crawlers each site blocks. Whether it publishes llms.txt or agents.md. Whether it quietly serves crawlers something different from what it serves you. And which platforms and CDNs are making that decision on the operator's behalf.

See the indexCheck any domain

What the last crawl found

3,678 domains measured on 2026-08-09. Sites we could not reach are excluded rather than counted as failures.

18%
Block at least one answer-surface crawler
661 of 3,678 sites
12%
Publish an llms.txt
448 sites
26%
Refuse GPTBot at the server
953 sites, whatever robots.txt says
70
Mean readiness score
Out of 100, across scored sites

Who is actually deciding

Most operators never formed a policy on AI crawlers. Their edge network did, by default, and they inherited it. Sites behind fastly block an answer-surface crawler 24.8% of the time, against 3.2% behind google.

AI blocking rate by edge network
Edge networkSitesBlocking AIProportion blockingMean score
Cloudflare1,076188 (17.5%)71.7
Amazon CloudFront42487 (20.5%)63.6
Akamai30744 (14.3%)63.8
Fastly22255 (24.8%)65.9
Google Cloud / GFE1545 (3.2%)61
Vercel483 (6.3%)76.5

All edge networks . By publishing platform

Readiness by platform

What a site is built on predicts how legible it is to an agent.

AI blocking rate by publishing platform
PlatformSitesBlocking AIProportion blockingMean score
WordPress29852 (17.4%)72.3
Next.js29548 (16.3%)67.3
Adobe Experience Manager1566 (3.8%)67.6
Drupal1204 (3.3%)64.3
HubSpot CMS783 (3.8%)74.2
Nuxt574 (7.0%)63.6

Which crawlers get shut out

Share of measured sites whose robots.txt denies each crawler the site root.

AI crawlers ranked by how many indexed sites block them
CrawlerOperatorBlocked byProportion
CCBotCommon Crawl585 (15.9%)
GPTBotOpenAI549 (14.9%)
BytespiderByteDance548 (14.9%)
ClaudeBotAnthropic529 (14.4%)
meta-externalagentMeta493 (13.4%)
Google-ExtendedGoogle478 (13.0%)
Applebot-ExtendedApple456 (12.4%)
AmazonbotAmazon445 (12.1%)
DiffbotDiffbot387 (10.5%)
cohere-aiCohere386 (10.5%)
PerplexityBotPerplexity383 (10.4%)
YouBotYou.com331 (9.0%)

All 23 tracked crawlers

Least agent-ready right now

Fully measured sites with the lowest scores. Partial assessments are excluded because a renormalised score is not comparable to a complete one.

Lowest scoring domains in the index
RankDomainScoreAnswer-surface crawlersAgent files
1228furaffinity.netrefused GPTBotScore 11 out of 100, grade F11 blockednone
56tiktok.comScore 15 out of 100, grade F11 blockednone
501amazon.co.jpScore 15 out of 100, grade F10 blockednone
612amazon.frScore 15 out of 100, grade F10 blockednone
734amazon.esScore 15 out of 100, grade F10 blockednone
3508amazon.aeScore 15 out of 100, grade F10 blockednone
25amazon.comScore 18 out of 100, grade F10 blockednone
317amazon.co.ukScore 18 out of 100, grade F10 blockednone
382amazon.deScore 18 out of 100, grade F10 blockednone
642amazon.caScore 18 out of 100, grade F10 blockednone

Full leaderboard, best and worst

Why this exists

Publishers are deciding, one robots.txt at a time, whether AI systems may read the web. Those decisions are made quietly, changed without announcement, and are individually trivial to check but collectively invisible.

CrawlIndex checks them on a schedule and keeps the receipts. The rubric is published, every score is arithmetic over archived evidence, and no language model touches the numbers. The whole dataset is downloadable. If you disagree with a result you can read exactly how it was reached and recompute it yourself.

Read the methodology . Download the dataset

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "The state of AI crawler access." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.