
Claude Sonnet 4.6 Arrives With 1M Context and Near-Opus Coding Performance
Anthropic's new mid-tier model matches Opus 4.6 on coding benchmarks, ships a million-token context window, and keeps the same $3/$15 pricing as its predecessor.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Anthropic's new mid-tier model matches Opus 4.6 on coding benchmarks, ships a million-token context window, and keeps the same $3/$15 pricing as its predecessor.

Four UC San Diego researchers argue in Nature that current LLMs already constitute artificial general intelligence, igniting fierce debate across the AI community.

xAI previews Grok 4.20 with enhanced multimodal capabilities and further reduced hallucinations, building on Grok 4.1's success. The company also teases a 6 trillion parameter Grok 5.

Alibaba releases Qwen 3.5, a 397B parameter open-source multimodal model with 256K context, Apache 2.0 license, and performance that tops Python coding and math reasoning benchmarks.

OpenAI releases GPT-5.3-Codex, a frontier coding model that is 25% faster, sets new records on SWE-Bench Pro and Terminal-Bench 2.0, and was instrumental in creating itself.

Anthropic launches Claude Opus 4.6 featuring agent teams, adaptive thinking, 1M token context window, and state-of-the-art performance on Terminal-Bench 2.0 and Humanity's Last Exam.

Z.ai releases GLM-5, a 744B parameter open-source Mixture-of-Experts model purpose-built for agentic tasks, scoring 77.8% on SWE-bench Verified and 56.2% on Terminal-Bench 2.0.

OpenAI begins testing advertisements in ChatGPT for Free and Go tier users in the US, while Plus, Pro, Business, Enterprise, and Education plans remain ad-free.

Apple is reportedly planning to integrate OpenAI's GPT-5 into Apple Intelligence across iOS 26, iPadOS 26, and macOS Tahoe, with major implications for Siri and the broader Apple AI strategy.

DeepSeek releases V3.2 under MIT license with 671B MoE architecture, matching GPT-5 at one-tenth the cost and achieving gold-medal performance on IMO and IOI competitions.

Analysis of how the MMLU benchmark gap between open-source and proprietary AI narrowed from 17.5 to 0.3 percentage points in a single year, reshaping the industry landscape.

The AI agent market reached $7.6 billion in 2025 with 49.6% projected annual growth. Gartner confirms 40% of enterprises will have dedicated AI agent teams by end of 2026.