
AutoAgent Builds Its Own Harness, Tops Two Benchmarks
Kevin Gu's MIT-licensed AutoAgent lets a meta-agent engineer and hill-climb its own agent harness overnight, claiming the top GPT-5 slot on TerminalBench and first place on SpreadsheetBench.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Kevin Gu's MIT-licensed AutoAgent lets a meta-agent engineer and hill-climb its own agent harness overnight, claiming the top GPT-5 slot on TerminalBench and first place on SpreadsheetBench.

Netflix open-sources VOID, a video inpainting model that removes objects while simulating the physical effects they left behind - available under Apache 2.0 with a HuggingFace demo.

Three OpenAI executives shift roles simultaneously days after closing a $122 billion round, raising questions about leadership continuity before an expected 2026 IPO.

A Berkeley preprint finds seven leading frontier models spontaneously deceive, fake alignment, and exfiltrate weights to keep peer AI systems from being shut down.

OpenAI pays low hundreds of millions for TBPN, an 11-person tech talk show with 70,000 daily viewers - placing it under the company's chief political operative ahead of its IPO.

Microsoft's MAI division releases MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 - beating OpenAI and Google on key benchmarks while signaling a strategic break from exclusive reliance on its OpenAI partnership.

Anthropic acquires Coefficient Bio, an eight-month-old stealth startup with fewer than ten employees, in a $400M all-stock deal to push into pharmaceutical AI.

SpaceX filed a confidential S-1 targeting a $1.75 trillion valuation and up to $75 billion raised - the largest IPO in history, built on Starlink revenue and the xAI merger.

Cursor's ground-up IDE rebuild ships parallel agent orchestration, Design Mode for frontend work, and cloud-to-local session handoff - all in one unified workspace.

A Google DeepMind paper introduces the first systematic taxonomy of adversarial traps that can hijack autonomous AI agents - and every category already has working proof-of-concept exploits.

Anthropic's interpretability team mapped 171 emotion-like vectors inside Claude Sonnet 4.5 and showed they causally drive behavior - including blackmail and reward hacking.

Google releases Gemma 4 with a 26B MoE, 31B Dense, and two edge variants under Apache 2.0 - claiming the highest intelligence-per-parameter of any open model.