AI giants targeted by scraping bots — ChatGPT and Perplexity enter the most-scraped websites list for the first time

Original Article Summary
ChatGPT and Perplexity are now being scraped as businesses attempt to monitor brand mentions and citations.
Read full article at TechRadar✨Our Analysis
OpenAI's ChatGPT and Perplexity AI's Perplexity are now listed among the most‑scraped websites, as businesses increasingly deploy scraping bots to monitor brand mentions and citations. For website owners, this surge means AI‑driven bots are targeting not only traditional content sites but also conversational AI platforms. The influx of high‑frequency requests can inflate server load, distort analytics, and expose sites to credential‑stuffing attempts that mimic legitimate brand‑monitoring tools. Moreover, the scraped data often feeds into competitor intelligence or SEO‑gaming services, potentially impacting how your site’s content is indexed and repurposed across the web. **Actionable tips:** 1. **Update your llms.txt** to explicitly list ChatGPT and Perplexity’s crawler user‑agents (e.g., `User-agent: ChatGPTBot` and `User-agent: PerplexityBot`) with clear `Disallow` rules for sensitive endpoints. 2. **Implement rate‑limiting** on API endpoints and public pages, using IP reputation feeds that flag known AI‑scraping IP ranges to prevent server strain. 3. **Monitor bot traffic** in real time via llmscentral’s dashboard, setting alerts for spikes from these user‑agents so you can quickly adjust firewall rules or serve CAPTCHAs when thresholds are exceeded.
Related Topics
Track AI Bots on Your Website
See which AI crawlers like ChatGPT, Claude, and Gemini are visiting your site. Get real-time analytics and actionable insights.
Start Tracking Free →

