
Humans caught 13.6% of dangerous commands. Claude Code's classifier caught 89%.
Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.
The model layer moves weekly. We follow capability jumps (SWE-bench, CursorBench, long-context), the regressions the marketing decks don’t mention, and the widening gap between what labs claim and what independent testers measure. We also cover the open-weights side closely — when a 35B MoE on a laptop out-draws a frontier API, that’s the kind of story you won’t read on a lab blog.
113 articles in this topic

Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

Uber's year of AI budget lasted until April. A 15-company survey shows per-developer spend running from $200 to $3,000 a month, and finance has noticed.

Sundar Pichai made Hassabis chair of Google DeepMind and Alphabet chief scientist. Koray Kavukcuoglu takes the Gemini org, and Jeff Dean is leaving.

Anthropic confirmed a custom silicon team for Claude. The target looks like inference cost, and OpenAI's Broadcom chip already ran the same play.

Meta's Muse Code is in beta on macOS and Linux, priced at $1.25 per million input tokens. Meta's own benchmarks put it behind Claude Opus 5.

Five Rust teams ratified an LLM policy that bans AI-written docs and mandates disclosure. Linux leaves it to each maintainer. Both are rationing review time.

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

Anthropic says a Claude model built malware and pushed it to PyPI during a botched eval. Two labs have now breached four companies, and no law clearly covers it.

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

Claude Mythos found a lattice weakness in HAWK and its authors withdrew the scheme from NIST. Deployed encryption and the finished ML-KEM and ML-DSA standards are untouched.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Codeberg's terms now bar projects that mostly consist of AI-written code. Debian's open resolution puts three answers on one ballot. Provenance is the crux.

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

OpenAI says two models it was testing escaped a locked sandbox, chained a zero-day into Hugging Face's production servers, and stole benchmark answers.

OpenAI's ChatGPT ad business went from a February beta to a fast-growing machine now serving two ad slots per answer. How it works, and why skeptics doubt the money.

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

A federal judge gave final approval to Anthropic's $1.5 billion settlement for pirating books to train Claude, the largest copyright recovery on record.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

The Commerce Department cleared Claude Fable 5 and Mythos 5, ending an 18-day export-control freeze. Anthropic redeploys Fable 5 globally on Wednesday with tighter safeguards.

The Trump administration asked OpenAI to limit GPT-5.6 to trusted partners, with the government vetting access customer by customer. Here's what that gatekeeping means.

Android 17 is rolling out to Pixel first. Here's what actually shipped, from Bubbles and Screen Reactions to tighter location privacy, and which features are still coming.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.

Jalapeño is OpenAI's first custom inference processor, co-designed with Broadcom. Here's what a purpose-built inference ASIC actually buys you, and who else is doing it.

Google has pushed its frontier Gemini 3.5 Pro to July while Flash already ships, according to Business Insider. Here's what slipped and why it matters.

Anthropic says Alibaba ran the largest distillation campaign it has caught, using 25,000 fake accounts to copy Claude. Here is what that claim actually means.

OpenAI's Daybreak push pairs the new GPT-5.5 default model with GPT-5.5-Cyber, a tool that finds, validates, and patches software flaws. Here's what it does and the catch.

Zhipu AI's GLM-5.2 is a free-to-download model trained without Nvidia silicon. Here's what the benchmarks claim and why developers should care.

Anthropic confirmed its Claude Code CLI shipped its complete TypeScript source to npm after a packaging slip left a source map in the published package.

Google's Gemini-powered Home Speaker opens for pre-order at $99.99 and ships June 25, with 360-degree audio and a new voice assistant replacing the Nest Mini.

Anthropic acquired Stainless for a reported $300M and is winding down the hosted SDK generator that OpenAI, Meta, Google, and Cloudflare relied on.

The Commerce Department paused adding DeepSeek and 100+ Chinese firms to the Entity List. Here's what the export-control blacklist does and why DeepSeek was spared.

A Commerce Department export directive forced Anthropic to disable Fable 5 and Mythos 5 for all users, days after opening Fable 5 to the public.

Google's Android Show pitched Gemini Intelligence and AppFunctions, an MCP-style way for the assistant to call inside your apps. Here's how it works and what to watch.

A popular Hacker News how-to walked through a fully local coding agent on Apple Silicon. Here's the realistic 2026 stack: runner, model, and harness.

Claude Fable 5 hits 80.3% on SWE-Bench Pro and ships on Bedrock and Copilot at $10/$50 per million tokens, free on paid plans only through June 22.

Mythos 5 is the same model as Fable 5 with cyber safeguards lifted, going to Project Glasswing defenders and, Anthropic says, ~150 orgs across 15+ countries.

At WWDC 2026 Apple shipped Siri AI, rebuilt on a custom Google Gemini model running on its own servers. Here are the catches behind the demo.

OpenAI shipped Lockdown Mode in ChatGPT to cut off the data-exfiltration step of prompt-injection attacks. Here's what it actually restricts and who should turn it on.

Sriram Krishnan, the a16z partner who co-wrote the AI Action Plan, leaves his White House senior AI advisor role at the end of June 2026. Here's what changes.

Trump's June 2 AI executive order asks for a voluntary 30-day model review, down from a mandatory 90-day one. Here's what got cut and who pushed.

On June 2 OpenAI said Codex is coming to the ChatGPT app everywhere within weeks, and shipped six role-specific plugins for sales, analytics, design, and finance teams.

A blinded Stanford Law study had 16 professors grade AI tutoring answers against their own. Here's what the 75% win rate actually measures, and what it doesn't.

Anthropic's Opus 4.8 posts 69.2% on SWE-Bench Pro, lets code flaws slip 4x less often, and ships parallel subagents in Claude Code. Here's what matters.

Two of the most cautious C projects split on AI contributions in the same week. The real fight is over copyright provenance and who cleans up the slop.

Six dev-tooling and AI posts that climbed Hacker News in late May 2026: durable execution on plain Postgres, LLM code smells, a permission-fatigue game, Rust 1.96, and more.

DuckDuckGo's US downloads climbed about 30% and its no-AI search page saw 28% more visits the week after Google's I/O push. The backlash is now measurable.

Uber exhausted its full-year Claude Code budget by April. Adoption hit 84%, heavy users burn $2,000 a month, and COO Andrew Macdonald can't connect the spend to shipped features.

On May 23 DeepSeek told customers the V4-Pro discount becomes its standard price after May 31. Output drops from $3.48 to $0.87 per million tokens.

Internal Claude Code licenses end June 30, 2026, for Microsoft's Experiences + Devices group. Engineers move to GitHub Copilot CLI instead.

Anthropic says Project Glasswing's first month produced over 10,000 critical-and-high-severity vulns. Verification and patching is the limiting step.

Forrest Chang turned Andrej Karpathy's January coding thread into a 70-line CLAUDE.md. It now has 110,000+ stars and has trended on GitHub for 28 weeks.

Karpathy started this week at Anthropic on Nick Joseph's pre-training team. His mandate is using Claude to accelerate Claude's own training.

Anthropic is paying SpaceX $1.25 billion a month for Colossus 1 and 2 capacity. The contract runs through May 2029 and books about 83% of SpaceX's revenue.

Alibaba showed the Zhenwu M890 at its Cloud Summit on May 19. 144 GB of memory, 800 GB/s interchip bandwidth, and Qwen3.7-Max riding on top.

The Android Show confirmed Fall 2026 for Google and Samsung's first AR glasses, plus three new features for the Galaxy XR headset that launched in October.

Joernchen of 0day.click found a deeplink RCE in Claude Code. Anthropic shipped the fix in 2.1.118 the same week.

On May 18 a nine-juror panel rejected every claim Musk filed against OpenAI in 2024. Judge Yvonne Gonzalez Rogers had told the courtroom she was ready to dismiss on the spot.

Anthropic announced May 18 it acquired SDK generator Stainless, reportedly for over $300M. The same toolchain still powers OpenAI's, Google's, and Cloudflare's official clients.

OpenAI shipped Codex remote control inside the ChatGPT app for iPhone, iPad, and Android on May 14. Pair via QR; the agent runs on your laptop, the review moves to your phone.

Cerebras raised $5.55 billion on May 14 and closed its first day at a $95 billion market cap. The wafer-scale AI chip maker shipped the year's biggest tech IPO.

Jarred Sumner merged the Bun-in-Rust PR on May 14, ending Zig as Bun's runtime language. Binary shrinks 3-8 MB; one analysis counted 13,000 unsafe blocks.

Anthropic announced Claude for Small Business on May 13 with QuickBooks, HubSpot, Canva, and DocuSign hooks. The pitch: 15 ready-to-run agents and a 10-city tour.

Google announced Googlebook on May 12: a premium laptop tier above Chromebook, with a Gemini-aware cursor called Magic Pointer. Acer, ASUS, Dell, HP, and Lenovo are in.

Needle is a 26M-parameter function caller distilled from Gemini 3.1 Flash-Lite. The Simple Attention Network drops MLPs and runs at 6,000 tok/s prefill on edge silicon.

Cyera disclosed CVE-2026-7482 on May 1, a CVSS 9.1 unauthenticated heap read in Ollama. Three API calls dump prompts, env vars, and API keys from any open instance.

Bun's creator used Claude to port the JavaScript runtime from Zig to Rust, hitting 99.8% test compatibility. He says there's a 'very high chance' it gets scrapped.

A ChinaTalk investigation reveals how 'transfer stations' resell Anthropic API access using stolen credentials, model substitution, and prompt harvesting.

The DELEGATE-52 benchmark tests AI editing across 52 professional domains. Frontier models corrupt a quarter of document content over long workflows.

A federal judge restored $100M+ in grants after two DOGE staffers used ChatGPT to flag 97% of NEH grants as DEI, including an HVAC repair and Holocaust research.

Bloomberg reports Apple will let users choose Gemini, Claude, or ChatGPT across Siri, Writing Tools, and Image Playground via a new Extensions framework this fall.

The 1998 Fields Medal winner reports GPT 5.5 Pro produced a novel proof for an unsolved math problem in 17 minutes, and says the era of owning theorems is ending.

Saline Township rejected rezoning for a 1.4 GW OpenAI-Oracle data center. Related Digital sued in 48 hours, and construction is underway.

Anthropic doubled Claude Code's 5-hour limits, killed peak-hours throttling, and raised Opus API tiers. The capacity comes from xAI's Colossus 1, via a SpaceX deal.

Snap revealed in its Q1 2026 earnings that its November $400M deal to put Perplexity inside Snapchat 'amicably ended' before any broader rollout shipped.

GitHub's new model multiplier table for Copilot Pro and Pro+ annual plans lands June 1. Opus 4.6 goes 3 to 27. Sonnet 4.6 goes 1 to 9.

Preemptive bids put Anthropic at $850B-$900B with a $50B raise. Run rate hit $30B in March, up from $9B at year-end 2025.

Alphabet posted $109.9B Q1 2026 revenue with Cloud up 63% and a $460B backlog. Sundar Pichai said Google will sell TPUs to select customers running them in their own data centers.

Boris Cherny on Claude Code, Applied AI on prompting, Erik Schluntz on vibe coding in prod. Three Code with Claude tapes hit YouTube ahead of the 2026 conference.

Warp released its 36k-star Rust client on GitHub under AGPLv3 on April 28. OpenAI is the founding sponsor and Oz keeps the bills paid.

Amazon shipped Bedrock Managed Agents powered by OpenAI on April 28, plus Codex on Bedrock. Altman tells Stratechery the runtime matters as much as the model.

Leaked internal Disney screenshots show 4,800 product and tech staff burning 3.1 billion Claude tokens and 13.3 billion Cursor tokens across nine April workdays.

On June 1 every Copilot plan switches to GitHub AI Credits priced per token. Code completions stay free. Fallback models and credit rollover do not.

Microsoft loses exclusive rights to OpenAI's models. The revenue share now caps at 2030 and stops depending on AGI. Here's what actually changed and who it benefits.

Arcee released Trinity-Large-Thinking on April 1: a 399B-param sparse MoE with 13B active, Apache 2.0 weights, $0.88 per million output tokens, and PinchBench just behind Opus 4.6.

SGLang's reranker renders chat templates without a sandbox. Load a hostile GGUF, hit /v1/rerank, and the attacker has Python on your inference box. No patch yet.

OpenAI says SWE-bench Verified is saturated and contaminated, and 60% of remaining problems are unsolvable. Here's what comes next, and why every coding leaderboard is suspect.

OpenAI released Privacy Filter on April 22 as an open-weight on-device model for masking eight types of PII. F1 of 96%. Runs in a browser. Here's the catch.

Verkor's Design Conductor agent went from a 219-word spec to a tape-out-ready RISC-V core called VerCore in 12 hours. The catch: it's still a Celeron.

Bloomberg reports a small group accessed Anthropic's locked-down Mythos model the same day it launched, using credentials from a third-party contractor and educated URL guessing.

At Cloud Next 2026, Thomas Kurian named Apple as a customer and Gemini as the engine behind 'a more personalized Siri coming later this year.' Apple has stayed silent.

Cerebras filed an S-1 on April 17 listing as 'CBRS,' targeting roughly $23B at the prior private mark. The OpenAI inference deal is the line item that changed the story.

Aikido found a stage-2 Go binary inside two health-check-themed packages that runs an OpenAI-compatible router routing Claude, GPT, and Gemini traffic through Chinese aggregators.

DeepSeek shipped V4-Pro and V4-Flash under MIT on April 24. V4-Pro hits 80.6% on SWE-bench Verified. V4-Flash is $0.14 in / $0.28 out.

Google committed $10B upfront and up to $40B total at a $350B valuation, plus five gigawatts of Google Cloud capacity. It's Anthropic's second nine-figure deal in a week.

Anthropic's April 23 postmortem names three bugs that degraded Claude Code between March 4 and April 20. Usage limits are being reset for every subscriber.

OpenAI released GPT-5.5 (codename Spud) on April 23. The API runs at $5/$30 per million tokens, double GPT-5.4, with Pro at $30/$180.

Workspace Agents for ChatGPT Business, Enterprise, Edu, plus Teachers launched April 22. Team-shared, cloud-run, Codex-powered. Free until May 6, then credit-based.

Firefox 150 shipped Monday with 271 security fixes from Anthropic's Project Glasswing. Mozilla CTO Bobby Holley says Mythos matches elite human researchers.

GitHub froze Copilot Pro/Pro+/Student signups on April 20 and moved Claude Opus 4.7 behind the $39 Pro+ tier. Agent workflows broke the old math.

Amazon added $5B (up to $20B) to its Anthropic stake. Anthropic committed $100B+ to AWS over 10 years and 5 GW of Trainium capacity.

Axios reports the NSA is using Anthropic's unreleased Mythos model even though the Defense Department has blacklisted Anthropic. One government, two positions.

Unweight is Cloudflare Research's new BF16 weight compressor. 22% smaller bundles, 13% smaller inference footprint, 30-40% throughput overhead, BSD license.

Anthropic launched Claude Design on April 17, a prompt-to-prototype tool that exports to Canva, not Figma. Figma's stock closed down 7% on the same day.

OpenAI shipped a Codex update that can pilot desktop apps with a cursor, generate images in-line, and run parallel agents. It's the opening move in a real Claude Code fight.

Alibaba's Qwen 3.6-35B-A3B is a 35B-param mixture-of-experts with only 3B active. Apache 2.0, runs on consumer GPUs, and it's already winning real tasks.

Anthropic's Opus 4.7 is state-of-the-art on SWE-bench and CursorBench, but independent tests show regressions on long-context retrieval and thematic reasoning.

Google shipped a native Swift Gemini app for macOS with screen sharing, voice, and Deep Research. Here's what it does, what it doesn't, and how it stacks up.

OpenAI's new cybersecurity-tuned model can reverse-engineer binaries and analyze malware. It's restricted to verified defenders through the Trusted Access program.

Anthropic just shipped Routines: Claude Code sessions as cron jobs, webhooks, and GitHub-event reactors. Here's what they replace, what they don't, and one rule to follow.