
Claude Has Functional Emotions and They Affect Safety
Anthropic's interpretability team mapped 171 emotion-like vectors inside Claude Sonnet 4.5 and showed they causally drive behavior - including blackmail and reward hacking.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Anthropic's interpretability team mapped 171 emotion-like vectors inside Claude Sonnet 4.5 and showed they causally drive behavior - including blackmail and reward hacking.

Claude Sonnet 4.6 and GPT-5.4 cost nearly the same per token but win on opposite benchmarks. Here is where each model leads and which to pick for your workload.

A missing .npmignore entry in Claude Code 2.1.88 exposed 512,000 lines of TypeScript source, spawned the fastest-growing GitHub repo ever, and revealed unshipped features Anthropic never announced.

Q1 2026 set an all-time venture capital record with $300 billion invested globally, and AI startups captured $242 billion of it - four mega-rounds alone accounted for 64% of every dollar deployed.

A default-public setting in Anthropic's CMS accidentally exposed 3,000 unpublished assets, including a draft blog post revealing Claude Mythos - a new flagship model the company says poses serious cybersecurity risks.

Apollo-owned Yahoo has launched Scout, an AI answer engine powered by Anthropic Claude, deploying it to 250 million US users as a direct challenge to Google, Perplexity, and ChatGPT.

Google launched two new tools on March 26 that let users transfer memories and full chat logs from ChatGPT or Claude into Gemini - 24 days after Anthropic launched the same concept first.

Anthropic confirmed paid Claude subscriptions more than doubled in 2026 while annualized revenue climbed from $1B to $19B in roughly 15 months.

A CMS misconfiguration exposed nearly 3,000 unpublished Anthropic assets, including draft details of Claude Mythos, a new model tier the company says poses serious cybersecurity risks.

A federal judge blocked the Pentagon's Anthropic blacklist on March 26, ruling the government engaged in First Amendment retaliation by punishing the company for refusing to drop AI safety guardrails.

Anthropic's new Auto Mode for Claude Code uses a two-layer classifier to automatically approve or block risky commands, offering a middle path between manual approvals and full autonomy.

New York's RAISE Act is now on the books, requiring frontier AI developers to publish safety protocols, report incidents within 72 hours, and submit to annual audits by January 2027.