Recent Articles - Page 11

Latest News

Anthropic's $1.5B Book Piracy Settlement Wins Approval

Anthropic's $1.5B Book Piracy Settlement Wins Approval

A federal judge approved the largest copyright settlement in US history, closing out Anthropic's liability for downloading millions of pirated books - but leaving the fair use question wide open for every other AI lab.

View All News →

Guides

View All →

Reviews

View All →
Kimi K3 Review: Best at Code, Worse at Honesty

Kimi K3 Review: Best at Code, Worse at Honesty

Moonshot's Kimi K3 tops LMArena's Frontend Code Arena and undercuts Opus 4.8 on cost per task, but a tripled price tag, a rising hallucination rate, and an unresolved distillation question complicate the win.

Leaderboards

View All →

Models

View All →
Qwen3-VL-235B-A22B

Qwen3-VL-235B-A22B

Alibaba's flagship open-weight vision-language MoE beats every proprietary model on DocVQA at 96.5% and MathVista at 85.8%, but trails GPT-5.4 and Gemini 3.1 Pro on broad MMMU-Pro reasoning.

DeepSeek-VL2

DeepSeek-VL2

DeepSeek-VL2 is DeepSeek's open-weight Mixture-of-Experts vision-language model, activating just 4.5B of its 27B parameters to hit 93.3% on DocVQA and beat GPT-4o on OCRBench.

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL-72B-Instruct

Alibaba's dense 72B vision-language model tops the open-weight DocVQA leaderboard at 96.4% and remains the default self-hosted choice for document and chart understanding.

Recent

JadePuffer Shows How AI Agents Now Run Ransomware

JadePuffer Shows How AI Agents Now Run Ransomware

Sysdig documents the first AI-agent ransomware operation: an LLM exploited CVE-2025-3248 in Langflow, moved laterally, and encrypted 1,342 production database records with no human directing each step.

Grok 4.1 Fast

Grok 4.1 Fast

Grok 4.1 Fast is xAI's agent-optimized model with a 2M-token context window, #1 ranking on tau-bench Telecom, and one of the lowest input prices among frontier-adjacent APIs at $0.20/M tokens.

GPT-5.4 mini

GPT-5.4 mini

OpenAI's mid-range model in the GPT-5.4 family delivers near-flagship coding and agentic performance at $0.75/M input tokens with a 400K context window.

LongCat-2.0 Review: China's Stealth Coder

LongCat-2.0 Review: China's Stealth Coder

Meituan's 1.6T open-source coding model secretly topped OpenRouter for two months before revealing itself - and the price-to-performance math is hard to argue with.