
Meta MTIA 400: GenAI Inference at Scale
Meta's second-gen ASIC delivers 6 PFLOPS FP8 and 288 GB HBM for GenAI and recommendation inference inside Meta's data centers.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Meta's second-gen ASIC delivers 6 PFLOPS FP8 and 288 GB HBM for GenAI and recommendation inference inside Meta's data centers.

Cisco reported record quarterly revenue of $15.8B and immediately announced 4,000 layoffs, raising its FY2026 AI infrastructure order target from $5B to $9B.

Bloomberg reports Google is in talks with SpaceX to launch its Project Suncatcher satellites - TPU-equipped spacecraft designed to run ML workloads in low Earth orbit.

Baiju Bhatt's orbital AI compute startup raises $275M at a $2B valuation and plans to build its own rockets to bypass SpaceX and Blue Origin bottlenecks.

AI2's federally backed OMAI compute cluster is now running on NVIDIA Blackwell Ultra hardware and has already shipped OLMo, Molmo 2, and MolmoAct models fully open to researchers.

NVIDIA and IREN plan 5 GW of DSX-aligned AI factories, backed by a $2.1B investment warrant and a $3.4B, five-year GPU cloud contract.

Anthropic gains 220,000 GPUs from SpaceX's Colossus 1 in Memphis, immediately doubling Claude Code five-hour rate limits for all paid plans.

Meta posted a record Q1 2026 revenue of $56.3 billion on April 29, then announced 8,000 layoffs and raised its AI infrastructure budget to $145 billion - sending the stock down 7% despite the record earnings.

Six companies just released MRC, an open networking protocol that routes AI training traffic across hundreds of simultaneous paths to end GPU idle time at supercomputer scale.

Anthropic has committed $200 billion to Google Cloud over five years - the largest cloud contract in AI history - alongside a 3.5 GW TPU capacity deal with Google and Broadcom coming online in 2027.

Bernstein projects Nvidia's China AI chip share falls from 66% to 8% in 2026 while Huawei targets $12B in revenue, with ByteDance alone committing $5.6B in Ascend 950PR orders.

Google's TPU 8i is a purpose-built inference chip with 10.1 FP4 PFLOPs, 288GB HBM3e at 8,601 GB/s, and a Boardfly topology that cuts collective latency 5x for agentic AI workloads.