Show HN: Directory of 28 AI crawlers – robots.txt rules, IP ranges, open data
Original Article Summary
Directory of all 28 AI crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more — with robots.txt rules to allow or block each one.
Read full article at Geoprompttracker.com✨Our Analysis
Geoprompttracker's release of a Directory of 28 AI crawlers—including GPTBot, ClaudeBot, PerplexityBot, and Google‑Extended—with detailed robots.txt rules and IP ranges gives website owners a concrete reference for managing automated AI traffic. For site operators, this means you can now differentiate between benign indexing bots and those that scrape content for generative models. The list reveals which crawlers respect standard robots.txt directives and which require explicit IP whitelisting or blocking. Ignoring these nuances could lead to unintended data exposure, higher server load, or compliance issues if AI models repurpose proprietary content. Actionable steps: 1) Update your llms.txt file to explicitly list the 28 crawler user‑agents and their allowed/disallowed paths, mirroring the rules in the directory. 2) Implement IP‑based firewall rules using the provided IP ranges for bots that ignore robots.txt, ensuring only vetted AI crawlers reach your server. 3) Use llmscentral’s monitoring dashboard to track hits from these agents in real time, setting alerts for any surge that deviates from the documented behavior.
Related Topics
Track AI Bots on Your Website
See which AI crawlers like ChatGPT, Claude, and Gemini are visiting your site. Get real-time analytics and actionable insights.
Start Tracking Free →Related Articles

Chinese military researchers and tech giants caught using Claude — US frontier model, coded 16 air-defense suppression tools targeting Taiwan, drafted anti-torpedo specs, and fed 151 million training queries to Alibaba
9/13/2026
Anthropic boss calls for AI development to slow down
9/12/2026

Show HN: MCP that gives Codex/Claude your SEO and AI visibility data
9/12/2026
