
GPT-5.6
OpenAI's GPT-5.6 family - Sol, Terra, and Luna - sets a new Terminal-Bench 2.1 record at 91.9% with subagent Ultra mode, but remains locked to ~20 government-vetted partners as of launch.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

OpenAI's GPT-5.6 family - Sol, Terra, and Luna - sets a new Terminal-Bench 2.1 record at 91.9% with subagent Ultra mode, but remains locked to ~20 government-vetted partners as of launch.

A 57-page DeepMind paper by co-founder Shane Legg identifies four pathways from AGI to superintelligence and six bottlenecks that could block each route.

Claude Mythos 5 is the full release of Anthropic's restricted Mythos family - same weights as Fable 5 but without safety classifiers for cybersecurity and biology, at $10/M input and $50/M output tokens.

Three new papers reveal how LLM safety hinges on persona training, how prompt modules interfere in deployed agents, and why scaling alone cannot reach symbolic reasoning.

The intelligence agencies of five allied nations issued a joint statement warning that frontier AI will fundamentally transform offensive cybersecurity within months, not years - and that most organizations are not ready.

The Trump administration is requiring OpenAI to vet every GPT-5.6 customer individually before granting access, citing cybersecurity capabilities that rival Anthropic's restricted Mythos model.

Three papers from today's arXiv: a 32B medical model beats DeepSeek-R1 in rare disease diagnosis, a KV cache method keeps 97% accuracy with 3% memory, and a new benchmark red-teams agentic AI systems.

OpenAI's GPT-5.5-Cyber found CVE-2026-8390 in Firefox's WebAssembly engine before Pwn2Own Berlin - five of six registered exploit entries withdrew.

The White House won't lift its ban on Anthropic's Fable 5 until the model can be made jailbreak-proof. Security experts explain why that condition is technically impossible.

New reporting reveals Amazon CEO Andy Jassy flagged a Fable 5 jailbreak on a routine White House call, triggering a 90-minute ultimatum that shut down Anthropic's two best models worldwide.

Three arXiv papers: a conscience mechanism for ethical training, shared memory for agent populations, and selective verification that cuts test-time compute waste.

Pramaana Labs uses the LEAN proof language to attach a mathematical certificate to every AI answer in high-stakes domains like tax, law, and drug discovery.