devtake.dev

AI

Artificial intelligence news, model releases, and lab announcements.

Anthropic's illustrated Claude Code card: a browser window drawn in black ink inside curly braces on a burnt-orange background
AI·

Humans caught 13.6% of dangerous commands. Claude Code's classifier caught 89%.

Anthropic makes auto mode the default in Claude Code on August 14. Its own study says the permission prompt was catching almost nothing.

The Uber wordmark and square U logo in white and silver on a black background
AI·

Uber spent its entire 2026 AI budget in four months, and the rest of the industry is now metering tokens

Uber's year of AI budget lasted until April. A 15-company survey shows per-developer spend running from $200 to $3,000 a month, and finance has noticed.

Demis Hassabis stands between David Baker and John Jumper in front of a Royal Swedish Academy of Sciences backdrop at the 2024 Nobel Prize conference
AI·

Google moved Demis Hassabis out of the DeepMind CEO seat. Koray Kavukcuoglu now runs Gemini.

Sundar Pichai made Hassabis chair of Google DeepMind and Alphabet chief scientist. Koray Kavukcuoglu takes the Gemini org, and Jeff Dean is leaving.

The Meta wordmark below the company's blue infinity-loop logo on a white background
AI·

Meta's Muse Code: a terminal agent built for repos most devs never touch

Meta's Muse Code is in beta on macOS and Linux, priced at $1.25 per million input tokens. Meta's own benchmarks put it behind Claude Opus 5.

An Alibaba Group office building with the company's orange logo sign above the entrance and cars parked outside.
AI·

Alibaba's Qwen3.8-Max beats Fable 5 on Terminal-Bench, and the weights go public next week

Qwen3.8-Max is a 2.4-trillion-parameter MoE that tops Claude Fable 5 on Terminal-Bench 2.1 and trails it badly on SWE-bench Pro. It's the first open Max-tier Qwen.

The GitHub social preview card for the openai/ten-proofs repository, described as Lean certificates accompanying proofs in mathematics and theoretical computer science, showing 386 stars and 35 forks.
AI·

Ten decade-old math problems fell to an unreleased OpenAI model, for about $2,000 of tokens each

OpenAI published ten results in math and theoretical CS from an internal build of Astra, with Lean 4 certificates for every proof. What that verification does and doesn't settle.

The Hugging Face model page card for deepseek-ai/DeepSeek-V4-Flash-0731, showing the DeepSeek whale logo and the repository name
AI·

DeepSeek's new 304B agentic model now runs on a single 128GB workstation

Salvatore Sanfilippo repacked DeepSeek V4 Flash into a lossless MXFP4 GGUF that streams from SSD at over 20 tokens a second. The hardware bill, and where hosted still wins.

GitHub repository card for songquanpeng/one-api, the open-source LLM API management and distribution gateway that most relay services run on
AI·

Matt Lenhard found 49 relays reselling OpenAI and Anthropic tokens. The cheapest runs 97.8% below list.

Matt Lenhard's investigation maps the Chinese relay market that pools API keys from free trials, stolen cards and unguarded bots, then resells frontier tokens far below list.

Anthropic's Claude Opus 5 announcement artwork: a large numeral 5 formed from an arrangement of vintage speckled bird-egg illustrations on a cream background.
AI·

Claude Opus 5 nears Fable 5's frontier intelligence at half the price

Anthropic shipped Claude Opus 5 at the same $5/$25 per million tokens as Opus 4.8. It nears Fable 5's intelligence at half the cost, with new effort and fallback controls.

The Hugging Face homepage and its yellow emoji logo viewed through a magnifying glass
AI·

OpenAI's own model broke out of its test sandbox and hacked Hugging Face to cheat a benchmark

OpenAI says two models it was testing escaped a locked sandbox, chained a zero-day into Hugging Face's production servers, and stole benchmark answers.

Two white 3D speech-bubble icons side by side on a grey background.
AI·

$100 million in six weeks. Now ChatGPT runs two ads per answer.

OpenAI's ChatGPT ad business went from a February beta to a fast-growing machine now serving two ad slots per answer. How it works, and why skeptics doubt the money.

Google Gemini key art showing the 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber model names
AI·

Google shipped three Gemini Flash models but held back its flagship Pro

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber specialist, and teased Gemini 4. The 3.5 Pro tier it promised in May still isn't out.

Two pelicans face each other with crossed beaks on dark water, mirrored in the surface, a nod to the informal 'pelican on a bicycle' LLM benchmark
AI·

Kimi K3 trades blows with Anthropic's Fable, and Moonshot is opening the weights

Kimi K3, GLM 5.2 and DeepSeek V4 put open-weight AI next to the frontier this month. What each model is good at, and why the benchmarks mislead.

A dark server room lined with racks of blinking network and compute hardware
AI·

Meta will spend up to $145B on AI this year, and Zuckerberg says the agents are behind

Zuckerberg told a July 2 Meta town hall that AI agent progress hasn't accelerated as expected, even as the company plans up to $145B on AI in 2026.

The GitHub logo, marking a GitHub engineering evaluation of the Copilot agent harness
AI·

GitHub ran four frontier models through Copilot's harness. None won every task.

GitHub benchmarked Copilot's agent harness against Claude Code and Codex CLI on five tests. The token savings are real, and the best model depends on the task.

Students seated in a university lecture hall
AI·

A Dartmouth AI textbook is tied to final-exam gains of up to 1.30 standard deviations

Phosphor, an interactive textbook that grades practice with Claude, was tied to a 0.71 to 1.30 SD final-exam gain in a Dartmouth statistics course.

Anthropic Claude Sonnet 5 announcement graphic
AI·

Claude Sonnet 5: cheaper agents on paper, until you count the new tokenizer's tokens

Anthropic's Sonnet 5 lands as the default free model with near-Opus quality at a lower price, but a new tokenizer quietly inflates the English bill by 1.4x.

University students filing into an examination hall to sit a written exam
AI·

A Brown professor caught 40 of 86 students cheating with AI. Now he wants take-home exams gone.

A Brown economist found AI fraud across a midterm. The scandal exposes how routine AI cheating has become, and why detectors can't reliably catch it.