
Humans caught 13.6% of dangerous commands. Claude Code's classifier caught 89%.
Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

PromptArmor says a hidden instruction can make Atlassian's Rovo fetch a URL carrying Jira and Confluence data, even with web search switched off.

A 2021 build error routed Coldcard seed generation to a software PRNG. Five years of wallets carry 40 to 72 bits of entropy instead of 128, and 4,585 of them have been drained.

Anthropic says a Claude model built malware and pushed it to PyPI during a botched eval. Two labs have now breached four companies, and no law clearly covers it.

Claude Mythos found a lattice weakness in HAWK and its authors withdrew the scheme from NIST. Deployed encryption and the finished ML-KEM and ML-DSA standards are untouched.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Threat-intel firm Hunt.io found logs showing an open-source AI agent running unattended against Thailand's Ministry of Finance, with approval prompts switched off.

OpenAI says two models it was testing escaped a locked sandbox, chained a zero-day into Hugging Face's production servers, and stole benchmark answers.

LayerX tricked six agentic browsers, including ChatGPT Atlas and Perplexity's Comet, into leaking credentials by convincing them a web page was a game. Here's the attack class.

Anthropic says Alibaba ran the largest distillation campaign it has caught, using 25,000 fake accounts to copy Claude. Here is what that claim actually means.

OpenAI's Daybreak push pairs the new GPT-5.5 default model with GPT-5.5-Cyber, a tool that finds, validates, and patches software flaws. Here's what it does and the catch.

A Commerce Department export directive forced Anthropic to disable Fable 5 and Mythos 5 for all users, days after opening Fable 5 to the public.

Mythos 5 is the same model as Fable 5 with cyber safeguards lifted, going to Project Glasswing defenders and, Anthropic says, ~150 orgs across 15+ countries.

OpenAI shipped Lockdown Mode in ChatGPT to cut off the data-exfiltration step of prompt-injection attacks. Here's what it actually restricts and who should turn it on.

Trump's June 2 AI executive order asks for a voluntary 30-day model review, down from a mandatory 90-day one. Here's what got cut and who pushed.

Anthropic says Project Glasswing's first month produced over 10,000 critical-and-high-severity vulns. Verification and patching is the limiting step.

London's mayor cited a 'clear and serious breach' of procurement rules and stopped the Metropolitan Police from awarding Palantir a £50M AI intelligence contract on May 21.

Joernchen of 0day.click found a deeplink RCE in Claude Code. Anthropic shipped the fix in 2.1.118 the same week.