Technical Specification & Research Standards
AI Crawler Access Index Methodology
This document describes the exact architecture, domain sampling list, HTTP fetch policy, parsing logic, and technical limitations of the SearchLearningTools AI Crawler Access Index.
1. Domain Sampling Source
Domains are sourced from the Tranco top web domains list (List ID: N9X52). Tranco provides a research-grade aggregate ranking of top domains derived from multiple public DNS and traffic metrics. To ensure reproducible benchmarking, the domain sample list ID is recorded with every dataset snapshot.
2. Verified AI Crawler User-Agents
We track 11 verified AI crawler and fetcher User-Agent tokens. Every User-Agent token in this index is verified directly against official vendor documentation. Unverifiable or deprecated tokens are omitted.
| User-Agent Token | Operator | Purpose | Official Documentation |
|---|---|---|---|
| GPTBot | OpenAI | AI Model Training | Docs |
| OAI-SearchBot | OpenAI | AI Search Retrieval | Docs |
| ChatGPT-User | OpenAI | Direct User Prompt Fetching | Docs |
| ClaudeBot | Anthropic | AI Model Training & Retrieval | Docs |
| PerplexityBot | Perplexity AI | AI Answer Retrieval | Docs |
| Google-Extended | Gemini AI Model Training | Docs | |
| CCBot | Common Crawl | Open Web Dataset Crawling | Docs |
| Applebot-Extended | Apple | Apple Intelligence Training | Docs |
| Bytespider | ByteDance | AI Training & Indexing | Docs |
| Amazonbot | Amazon | Web Indexing & AI | Docs |
| Meta-ExternalAgent | Meta | Meta AI Model Training | Docs |
3. HTTP Fetch & Scanner Execution
The automated scanner script (scripts/scan-ai-crawlers.ts) operates under strict compliance rules:
- Path Sourced: Fetches only
/robots.txtover HTTPS (with HTTP fallback). - Request User-Agent:
SearchLearningToolsBot/1.0 (+https://searchlearningtools.com/bot) - Timeout: Strict 5-second timeout enforced via
AbortController. - Redirect Policy: Follows up to a maximum of 3 redirects.
- Payload Cap: Enforces a 500 KB response size limit to prevent memory exhaustion.
- Rate Limiting & Concurrency: Concurrency limited to 10 requests with exponential backoff on HTTP 429 response codes.
4. Rule Evaluation Engine
Robots.txt contents are evaluated using the open-source robots-parser engine. For each domain and bot User-Agent, the root path (/) is evaluated to determine one of five deterministic statuses:
No disallow rule applies to the root path for this User-Agent or wildcard (*).
An explicit Disallow: / rule applies to this User-Agent or wildcard.
Root path is allowed, but specific subpath disallow rules apply.
HTTP 404 response received; all bots are implicitly allowed access.
- Advisory Policy Only:
robots.txtis an advisory protocol for web crawlers. It does not constitute technical firewall enforcement, IP blocking, or network access control. - Not a Measure of Actual Traffic: A blocked status in
robots.txtindicates site owner policy preference; it does not measure whether an AI bot actually attempted to crawl or fetch the domain. - Dynamic & User-Agent Truncation: Some CDN firewalls modify
robots.txtresponses dynamically based on request headers or IP reputation.
6. Methodology Changelog
v1.0 (2026-10-02): Initial methodology release tracking 11 AI bot user-agents across Tranco top domains.