
AI·
Humans caught 13.6% of dangerous commands. Claude Code's classifier caught 89%.
Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

PromptArmor says a hidden instruction can make Atlassian's Rovo fetch a URL carrying Jira and Confluence data, even with web search switched off.

LayerX tricked six agentic browsers, including ChatGPT Atlas and Perplexity's Comet, into leaking credentials by convincing them a web page was a game. Here's the attack class.

OpenAI shipped Lockdown Mode in ChatGPT to cut off the data-exfiltration step of prompt-injection attacks. Here's what it actually restricts and who should turn it on.