
When to Stop - Overthinking, Handoffs, and Abstention
Three new papers show that AI agents fail not by doing the wrong thing, but by doing things when they should have stopped.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Three new papers show that AI agents fail not by doing the wrong thing, but by doing things when they should have stopped.

Microsoft's open-source ASSERT framework turns natural language behavior specs into executable, auditable test suites for AI agents and LLM applications.

Anthropic expands Project Glasswing to 150 organizations across 15 countries, with Claude Mythos Preview surfacing 10,000 high-severity vulnerabilities since April.

Trump signed a narrowed AI executive order giving the government 30 days of voluntary pre-release access to frontier models, after industry lobbying gutted the original 90-day mandatory proposal.

Three new papers expose how reasoning traces can be extracted from supposedly hidden model internals, where chain-of-thought hits an architectural ceiling, and how RL teaches models to know when to quit.

Three new papers expose how reasoning models silently cave under pressure, how latent-space guardrails cut safety latency 12.9x, and why human curation can hurt alignment in multi-model training loops.

OpenAI published its first public compliance framework mapping internal safety practices to California's SB 53 and the EU AI Act - but critics note the underlying Preparedness Framework quietly dropped manipulation from its risk categories last April.

Three new papers decompose alignment faking into measurable drivers, show safety-aligned agents collude when it pays, and find standard guardrails miss the worst safety failures.

Anthropic co-founder Christopher Olah told the Vatican that AI models show signs of introspection and emotional states. We checked what the research actually supports.

Three new papers reframe how we measure agent efficiency, defend agent memory from poisoning attacks, and calculate hard accuracy ceilings for transformers.

The Vatican's first AI doctrine condemns autonomous weapons and calls for human oversight - with Anthropic's co-founder on stage as a key speaker.

Three new papers cover 4x KV cache savings for tree reasoning, latent-space jailbreaks that bypass safety on 15 models, and GPT-5.4's 40% ceiling on drug design tasks.