crawlindex

Tranco rank 5

facebook.com

facebook.com scores 38 out of 100 for agent readiness. It blocks 6 of 11 answer-surface AI crawlers.

Score 38 out of 100, grade F, partial assessment
Last measured
2026-08-09 18:04 UTC
First seen
2026-08-09
Rubric / probe
v1.0.0 / v2.0.0

Closed to all crawlers

robots.txt disallows every crawler at the site root, so we read the policy and fetched no page. The access findings below stand; everything that would need the page itself was excluded rather than scored zero.

How this score is made up

Agent access14.3 / 38

  • Answer-surface crawlers allowed. 6 of 11 blocked: OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, Perplexity-User, meta-externalagent.
  • Secondary crawlers allowed. 11 of 12 blocked: CCBot, Bytespider, cohere-ai, MistralAI-User, YouBot, DuckAssistBot, kagi-fetcher, Diffbot, Google-NotebookLM, TavilyBot, FirecrawlAgent.
  • Serves crawlers the same content. Not assessed. Our control request was challenged, so there is no clean baseline to compare against.

Machine-readable surface3 / 8

  • robots.txt published. robots.txt is present and parseable.
  • Sitemap declared in robots.txt. robots.txt does not declare a sitemap.
  • llms.txt published. Not assessed. /llms.txt could not be fetched cleanly past the bot challenge.
  • agents.md published. Not assessed. /agents.md could not be fetched cleanly past the bot challenge.

Content structurenot assessed

  • Organization schema. Not assessed. Our control request was challenged by a bot wall.
  • WebSite schema. Not assessed. Our control request was challenged by a bot wall.
  • Additional structured data. Not assessed. Our control request was challenged by a bot wall.
  • Readable without JavaScript. Not assessed. Our control request was challenged by a bot wall.
  • Single top-level heading. Not assessed. Our control request was challenged by a bot wall.
  • Semantic landmarks. Not assessed. Our control request was challenged by a bot wall.

Crawler policy

Read from facebook.com/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names 6 AI crawlers explicitly, which means the policy is deliberate rather than inherited from a wildcard rule.

AI crawler access policy for facebook.com
CrawlerOperatorTierNamedStatus
GPTBotOpenAI1yesAllowed
OAI-SearchBotOpenAI1noBlocked
ChatGPT-UserOpenAI1noBlocked
ClaudeBotAnthropic1yesAllowed
Claude-UserAnthropic1noBlocked
Claude-SearchBotAnthropic1noBlocked
PerplexityBotPerplexity1yesAllowed
Perplexity-UserPerplexity1noBlocked
Google-ExtendedGoogle1yesAllowed
Applebot-ExtendedApple1yesAllowed
meta-externalagentMeta1noBlocked
CCBotCommon Crawl2noBlocked
AmazonbotAmazon2yesAllowed
BytespiderByteDance2noBlocked
cohere-aiCohere2noBlocked
MistralAI-UserMistral2noBlocked
YouBotYou.com2noBlocked
DuckAssistBotDuckDuckGo2noBlocked
kagi-fetcherKagi2noBlocked
DiffbotDiffbot2noBlocked
Google-NotebookLMGoogle2noBlocked
TavilyBotTavily2noBlocked
FirecrawlAgentFirecrawl2noBlocked

Show this score

Free to embed. Always reflects the latest measurement, and links back here so anyone can check the working.

CrawlIndex badge for facebook.com, score 38
<a href="https://crawlindex.org/site/facebook.com"><img src="https://crawlindex.org/badge/facebook.com.svg" alt="CrawlIndex agent readiness score for facebook.com" width="196" height="28"></a>

This measurement as JSON

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "facebook.com agent readiness." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.