
AI Models Pass Vision Tests Without Seeing the Images
A Stanford study shows frontier AI models achieve 70-80% of visual benchmark scores with no images provided, exposing a fundamental flaw in how multimodal AI is evaluated.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

A Stanford study shows frontier AI models achieve 70-80% of visual benchmark scores with no images provided, exposing a fundamental flaw in how multimodal AI is evaluated.

Gemini 2.5 Flash costs 10x less and runs 4x faster than Claude Sonnet 4.6, but trails badly on coding benchmarks - here is the full breakdown.

Rankings of AI models on IFEval and IFBench, the two main benchmarks for measuring how reliably LLMs follow precise formatting, length, and content constraints.

Father Brendan McGuire, a 60-year-old Silicon Valley priest and former tech executive, helped Anthropic rewrite the Claude Constitution after the company asked the Vatican for help because AI was moving too fast.

Anthropic released Claude Managed Agents in public beta today, a fully managed platform that handles sandboxing, state, and tool execution so developers can skip building agent infrastructure from scratch.

Anthropic's restricted Claude Mythos Preview model autonomously discovered thousands of high-severity vulnerabilities across every major OS and browser, including bugs hiding in plain sight for 27 years.

Project Glasswing unites AWS, Apple, Google, Microsoft, CrowdStrike, and seven other organizations with $100M in credits for Anthropic's restricted Mythos Preview model to patch critical infrastructure before attackers catch up.

Britain is offering Anthropic a £40M research lab and a dual London Stock Exchange listing after the Pentagon branded the AI company a supply-chain risk.

OpenAI, Anthropic, Google, and Microsoft are now sharing attack detection data through the Frontier Model Forum to collectively block Chinese adversarial distillation campaigns.

Anthropic's run-rate revenue has surpassed $30 billion in 2026, tripling from $9 billion at end of 2025, as the company secures 3.5 gigawatts of next-gen TPU compute from Broadcom starting in 2027.

The Justice Department is asking the Ninth Circuit to reverse the order that blocked the Pentagon's supply chain risk label on Anthropic and paused Trump's federal ban on Claude.

Anthropic acquires Coefficient Bio, an eight-month-old stealth startup with fewer than ten employees, in a $400M all-stock deal to push into pharmaceutical AI.