AI Crawler Access Index: Top Web Domains Dataset
Empirical research tracking web crawler permissions for AI model training and answer engines (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, etc.) across top web domains.
Answer-First Summary Metrics
Scan Date: 2026-10-02 (List: N9X52)As of 2026-10-02, analysis of robots.txt directives across the top 20 web domains reveals that 30.0% block OpenAI's GPTBot, 30.0% block Anthropic's ClaudeBot, and 30.0% block PerplexityBot. On average, top domains disallow 3.5 out of 11 AI crawlers.
Bot Access Breakdown
Comparison of block rates across all 11 tracked AI crawlers and user agents.
AI Crawler Block Rate Breakdown
Percentage of top 20 web domains blocking each AI bot via robots.txt
Full Domain Access Index
Filter and search permissions for top web domains (including SearchLearningTools.com).
| Rank | Domain | robots.txt | GPTBot | OAI-SearchBot | ChatGPT-User | ClaudeBot | PerplexityBot | Total Blocked |
|---|---|---|---|---|---|---|---|---|
| #1 | searchlearningtools.comThis Site | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #2 | google.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #3 | youtube.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #4 | facebook.com | 200 OK | Allowed | Blocked | Blocked | Allowed | Allowed | 5 / 11 |
| #5 | wikipedia.org | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #6 | yahoo.com | 200 OK | Blocked | Allowed | Blocked | Blocked | Blocked | 7 / 11 |
| #7 | reddit.com | 200 OK | Blocked | Blocked | Blocked | Blocked | Blocked | 11 / 11 |
| #8 | amazon.com | 200 OK | Blocked | Blocked | Blocked | Blocked | Blocked | 9 / 11 |
| #9 | twitter.com | 200 OK | Blocked | Blocked | Blocked | Blocked | Blocked | 11 / 11 |
| #10 | instagram.com | 200 OK | Blocked | Blocked | Blocked | Blocked | Blocked | 11 / 11 |
| #11 | linkedin.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #12 | bing.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #13 | github.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 1 / 11 |
| #14 | microsoft.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #15 | apple.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #16 | netflix.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 5 / 11 |
| #17 | wordpress.org | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #18 | cloudflare.com | 200 OK | Allowed | Allowed | Allowed | Allowed | Allowed | 0 / 11 |
| #19 | pinterest.com | 200 OK | Blocked | Blocked | Blocked | Blocked | Blocked | 11 / 11 |
| #20 | tumblr.com | Err | - | - | - | - | - | 0 / 11 |
Download Open Dataset
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). You are free to download, analyze, publish, and build upon this data with attribution to SearchLearningTools.com.
https://searchlearningtools.com/research/ai-crawler-index when citing or reproducing dataset figures.Check Your Site
Test your site's live robots.txt and see how your AI access rules compare to this index.
Read Methodology
Inspect sampling methods, parser logic, vendor bot docs, and robots.txt advisory disclaimers.
Press Briefing
Key findings, media graphics, and ready-to-publish analysis summaries for journalists.
Cite Dataset
Copy ready-made APA, MLA, and BibTeX citations for academic papers and research reports.