Ask a model if code is malicious and it reaches for its morals

Original Article Summary
Inside every model we measured, a question about malicious code runs on the machinery of moral judgment.
Read full article at Manifold.security✨Our Analysis
Manifold Security's research revealing that “a question about malicious code runs on the machinery of moral judgment” in every tested model highlights how AI systems are increasingly filtering requests based on perceived ethics. For website owners, this means AI crawlers and code‑generation bots that encounter security‑related prompts may self‑block or return sanitized responses, reducing the risk of automated code‑injection attempts but also potentially limiting legitimate security‑testing traffic. Sites that host developer tools, code snippets, or security forums should expect a dip in malicious‑code queries from AI agents, while still monitoring for bots that bypass moral filters by obfuscating intent. **Actionable tips:** 1. Update your llms.txt to explicitly allow benign code‑analysis bots (e.g., “GoogleBot‑CodeReview”) while disallowing “MalwareScanner‑AI” agents that claim to test malicious payloads. 2. Deploy a bot‑traffic analytics layer that flags requests containing code‑related keywords combined with moral‑judgment triggers (e.g., “should I run this exploit?”) to differentiate genuine developer queries from evasion attempts. 3. Regularly audit your server logs for AI‑generated “ethical” responses that may indicate a model’s internal filtering, and adjust your robots.txt or llms.txt entries accordingly to maintain the right balance between security and accessibility.
Track AI Bots on Your Website
See which AI crawlers like ChatGPT, Claude, and Gemini are visiting your site. Get real-time analytics and actionable insights.
Start Tracking Free →


