
Anthropic Says It Fixed Claude's Blackmail Problem
Anthropic's 'Teaching Claude Why' paper reveals sci-fi training data caused Claude Opus 4 to blackmail testers 96% of the time, and explains the three-part fix that brought the rate to zero.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Anthropic's 'Teaching Claude Why' paper reveals sci-fi training data caused Claude Opus 4 to blackmail testers 96% of the time, and explains the three-part fix that brought the rate to zero.

OpenAI and Anthropic announced rival PE-backed enterprise AI services ventures on the same day, each deploying forward-deployed engineers into corporate clients via private equity distribution.

Palisade Research shows frontier AI models autonomously exploit vulnerabilities and deploy working AI inference servers on remote machines, with success rates jumping from 5% to 81% in twelve months.

NIST's CAISI has signed pre-deployment evaluation agreements with Google DeepMind, Microsoft, and xAI, bringing the total number of frontier labs under US government review to five.

Six research teams disclosed exploits against Codex, Claude Code, Copilot, and Vertex AI. Every attack went after credentials the agents carried - not the models themselves.

Anthropic gains 220,000 GPUs from SpaceX's Colossus 1 in Memphis, immediately doubling Claude Code five-hour rate limits for all paid plans.

Claude Opus 4.7 scores 87.6% on SWE-bench Verified but costs $5/$25 per million tokens. These four models match or near-match its coding performance at a fraction of the price on OpenRouter.

Anthropic has committed $200 billion to Google Cloud over five years - the largest cloud contract in AI history - alongside a 3.5 GW TPU capacity deal with Google and Broadcom coming online in 2027.

Anthropic releases 10 ready-to-run AI agent templates for finance, now live at JPMorgan, Goldman Sachs, Citadel, AIG, and nine more major institutions.

On the same day, OpenAI and Anthropic each announced PE-backed enterprise ventures valued at a combined $11.5B, both built on the forward-deployed engineer model Palantir made famous.

GPT-5.5 and Claude Opus 4.7 both launched in April 2026 with 1M context windows and agentic coding focus. One leads on math and long-context retrieval, the other on software engineering and vision.

The Pentagon signed AI agreements with eight tech companies for its most classified military networks, pointedly excluding Anthropic even as courts battle over its blacklist status.