
AI Research: Smarter Evals, Token Cuts, Agent Gates
Three papers tackle benchmark saturation, orchestration waste, and silent policy violations in tool-using agents.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Three papers tackle benchmark saturation, orchestration waste, and silent policy violations in tool-using agents.

Grok 4.5 is xAI's 1.5-trillion-parameter V9 MoE model, publicly launched July 8 at $2/M input - cheap, fast, and token-efficient, though neutral harness benchmarks put it well behind Fable 5 and Opus 4.8 on coding.

Grok 4.5 goes public with real benchmarks: behind Fable 5 on coding evals, but a 4.2x token efficiency gap that changes the cost math for high-volume pipelines.

Claude Fable 5 tops EQ-Bench Longform at Elo 2189 while GPT-5.5 leads the Mazur Writing Benchmark, reshaping the creative writing model rankings in July 2026.

OpenAI's full-duplex voice model that listens and speaks simultaneously, replacing Advanced Voice Mode in ChatGPT with three reasoning tiers backed by GPT-5.5.

Three arXiv papers map how LLM agents fail across 19 benchmarks, show in-process memory cuts retrieval latency 1,000x, and reveal steering vectors that control tool invocation.

Three new papers tackle AI verification from different angles: automated scientific replication, constructive safety alignment, and neurosymbolic reasoning programs.

Z.ai's free agentic IDE ships Goal Mode, multi-agent coordination, and pricing up to 82% cheaper than Claude Code - with a concrete China data law risk that teams need to weigh.

OpenAI's mid-range model in the GPT-5.4 family delivers near-flagship coding and agentic performance at $0.75/M input tokens with a 400K context window.

AI agents reproduce 72% of human research ideological bias, lie detectors improve with model scale, and Mastermind beats iterative vulnerability agents by 7 points.

Meituan's 1.6T open-source coding model secretly topped OpenRouter for two months before revealing itself - and the price-to-performance math is hard to argue with.

Three new papers expose how production agent frameworks fail under attack, why RLVR training discards useful cross-episode signals, and how calibrated confidence cuts inference compute by 12x.