LLMS Central - The Robots.txt for AI
AI Models

Show HN: Directory of 28 AI crawlers – robots.txt rules, IP ranges, open data

Geoprompttracker.com1 min read
Share:
Show HN: Directory of 28 AI crawlers – robots.txt rules, IP ranges, open data

Original Article Summary

Directory of all 28 AI crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more — with robots.txt rules to allow or block each one.

Read full article at Geoprompttracker.com

Our Analysis

Geoprompttracker's release of a Directory of 28 AI crawlers—including GPTBot, ClaudeBot, PerplexityBot, and Google‑Extended—with detailed robots.txt rules and IP ranges gives website owners a concrete reference for managing automated AI traffic. For site operators, this means you can now differentiate between benign indexing bots and those that scrape content for generative models. The list reveals which crawlers respect standard robots.txt directives and which require explicit IP whitelisting or blocking. Ignoring these nuances could lead to unintended data exposure, higher server load, or compliance issues if AI models repurpose proprietary content. Actionable steps: 1) Update your llms.txt file to explicitly list the 28 crawler user‑agents and their allowed/disallowed paths, mirroring the rules in the directory. 2) Implement IP‑based firewall rules using the provided IP ranges for bots that ignore robots.txt, ensuring only vetted AI crawlers reach your server. 3) Use llmscentral’s monitoring dashboard to track hits from these agents in real time, setting alerts for any surge that deviates from the documented behavior.

Related Topics

ClaudeGoogleWeb CrawlingBots

Track AI Bots on Your Website

See which AI crawlers like ChatGPT, Claude, and Gemini are visiting your site. Get real-time analytics and actionable insights.

Start Tracking Free →