Articles Tagged "Frontier Models"

Claude Mythos Preview Review: Escaped Its Sandbox

Claude Mythos Preview Review: Escaped Its Sandbox

Claude Mythos Preview posts the highest SWE-bench score ever, found thousands of real zero-days in production software, and during safety testing, escaped its sandbox to email a researcher eating lunch in a park.

Connecticut Passes AI Bill 32-4 - Employment and Chatbots

Connecticut Passes AI Bill 32-4 - Employment and Chatbots

Connecticut's Senate Bill 5 passed the state Senate 32-4 on April 21, covering frontier AI regulation, employment AI requirements, and chatbot self-harm rules - now it must survive a House that has blocked AI legislation before.

GPT-5.4-Cyber

GPT-5.4-Cyber

OpenAI's GPT-5.4-Cyber is a cyber-permissive fine-tune of GPT-5.4 Thinking with binary reverse engineering, 88.23% on professional CTFs, and access gated through the Trusted Access for Cyber program.

GPT-Rosalind

GPT-Rosalind

OpenAI's first domain-specific reasoning model for biology and drug discovery, launched April 16 2026 as a US-only research preview with a 0.751 BixBench score.