Recent Articles - Page 6

Latest News

Anthropic's $1.5B Book Piracy Settlement Wins Approval

Anthropic's $1.5B Book Piracy Settlement Wins Approval

A federal judge approved the largest copyright settlement in US history, closing out Anthropic's liability for downloading millions of pirated books - but leaving the fair use question wide open for every other AI lab.

Google's Frozen v2 Chip Bakes Gemini Into Silicon

Google's Frozen v2 Chip Bakes Gemini Into Silicon

Google is reportedly building a chip line separate from its TPUs that hardwires parts of Gemini directly into silicon, promising up to 10x efficiency as a capacity crunch forces Cloud to turn away customers.

View All News →

Guides

View All →

Reviews

View All →
Kimi K3 Review: Best at Code, Worse at Honesty

Kimi K3 Review: Best at Code, Worse at Honesty

Moonshot's Kimi K3 tops LMArena's Frontend Code Arena and undercuts Opus 4.8 on cost per task, but a tripled price tag, a rising hallucination rate, and an unresolved distillation question complicate the win.

Leaderboards

View All →

Models

View All →
AlayaWorld

AlayaWorld

AlayaWorld is a 15B open-weight video diffusion world model from Alaya Lab that sustains interactive, camera-controllable environments past 60 seconds.

Gemini 3.6 Flash

Gemini 3.6 Flash

Google DeepMind's workhorse Flash model cuts output tokens 17% versus Gemini 3.5 Flash, drops output pricing to $7.50/M, and cuts DeepSWE task tokens by 65% while trailing GPT-5.6 Luna and Grok 4.5 on raw coding scores.

DeepSeek-R1

DeepSeek-R1

DeepSeek-R1 is the 671B-parameter open-weight reasoning model that matched OpenAI o1 on math and coding benchmarks and triggered a $589 billion single-day drop in Nvidia's market cap in January 2025.

Recent

Qwen3.6-Plus

Qwen3.6-Plus

Alibaba's 1M-token flagship agentic coding model posts 78.8% on SWE-bench Verified and undercuts Kimi K2.6 and Claude Opus on price, but ships with no weights and a mandatory reasoning tax.

Gemini 3 Pro

Gemini 3 Pro

Google DeepMind's Gemini 3 Pro debuted at 1501 Elo on LMArena with 91.9% on GPQA Diamond and a 1M-token context window, before Google retired it for Gemini 3.1 Pro.

Hermes 4.3

Hermes 4.3

Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained entirely on the Solana-secured Psyche network.

Nous Research Talks Put Open-Source Hermes at $1.5B

Nous Research Talks Put Open-Source Hermes at $1.5B

Nous Research is finalizing a round led by Robot Ventures and USV that would value the open-source Hermes agent maker at $1.5 billion, built on a training network that skips traditional data centers entirely.