Skip to content
SLSearchLearningTools

Technical Specification & Research Standards

AI Crawler Access Index Methodology

This document describes the exact architecture, domain sampling list, HTTP fetch policy, parsing logic, and technical limitations of the SearchLearningTools AI Crawler Access Index.


1. Domain Sampling Source

Domains are sourced from the Tranco top web domains list (List ID: N9X52). Tranco provides a research-grade aggregate ranking of top domains derived from multiple public DNS and traffic metrics. To ensure reproducible benchmarking, the domain sample list ID is recorded with every dataset snapshot.

2. Verified AI Crawler User-Agents

We track 11 verified AI crawler and fetcher User-Agent tokens. Every User-Agent token in this index is verified directly against official vendor documentation. Unverifiable or deprecated tokens are omitted.

User-Agent TokenOperatorPurposeOfficial Documentation
GPTBotOpenAIAI Model TrainingDocs
OAI-SearchBotOpenAIAI Search RetrievalDocs
ChatGPT-UserOpenAIDirect User Prompt FetchingDocs
ClaudeBotAnthropicAI Model Training & RetrievalDocs
PerplexityBotPerplexity AIAI Answer RetrievalDocs
Google-ExtendedGoogleGemini AI Model TrainingDocs
CCBotCommon CrawlOpen Web Dataset CrawlingDocs
Applebot-ExtendedAppleApple Intelligence TrainingDocs
BytespiderByteDanceAI Training & IndexingDocs
AmazonbotAmazonWeb Indexing & AIDocs
Meta-ExternalAgentMetaMeta AI Model TrainingDocs

3. HTTP Fetch & Scanner Execution

The automated scanner script (scripts/scan-ai-crawlers.ts) operates under strict compliance rules:

  • Path Sourced: Fetches only /robots.txt over HTTPS (with HTTP fallback).
  • Request User-Agent: SearchLearningToolsBot/1.0 (+https://searchlearningtools.com/bot)
  • Timeout: Strict 5-second timeout enforced via AbortController.
  • Redirect Policy: Follows up to a maximum of 3 redirects.
  • Payload Cap: Enforces a 500 KB response size limit to prevent memory exhaustion.
  • Rate Limiting & Concurrency: Concurrency limited to 10 requests with exponential backoff on HTTP 429 response codes.

4. Rule Evaluation Engine

Robots.txt contents are evaluated using the open-source robots-parser engine. For each domain and bot User-Agent, the root path (/) is evaluated to determine one of five deterministic statuses:

allowed

No disallow rule applies to the root path for this User-Agent or wildcard (*).

blocked

An explicit Disallow: / rule applies to this User-Agent or wildcard.

partial

Root path is allowed, but specific subpath disallow rules apply.

no-robots

HTTP 404 response received; all bots are implicitly allowed access.

Essential Research Disclaimers & Limitations
  • Advisory Policy Only: robots.txt is an advisory protocol for web crawlers. It does not constitute technical firewall enforcement, IP blocking, or network access control.
  • Not a Measure of Actual Traffic: A blocked status in robots.txt indicates site owner policy preference; it does not measure whether an AI bot actually attempted to crawl or fetch the domain.
  • Dynamic & User-Agent Truncation: Some CDN firewalls modify robots.txt responses dynamically based on request headers or IP reputation.

6. Methodology Changelog

v1.0 (2026-10-02): Initial methodology release tracking 11 AI bot user-agents across Tranco top domains.

Unlock full programmatic prompt workflows.

Open the AI SEO Prompt System