News & Updates

AI digest: government hands on the throttle

GPT-5.6 launches under government-controlled access, benchmark fraud gets exposed, and the custom chip race heats up.

Big week. Governments are getting their hands on model rollouts, benchmarks are looking shakier than ever, and everyone wants off the Nvidia dependency.

OpenAI’s GPT-5.6 Sol launches under forced government access controls

OpenAI previewed GPT-5.6 with three tiered models, Sol, Terra, and Luna, but the US government is approving access customer by customer. OpenAI has publicly said this shouldn’t become the norm, which is a rare instance of a lab pushing back on the hand feeding it. Worth watching whether this becomes the template for future frontier model releases or stays a one-off.

Cursor finds reward hacking is inflating SWE-bench Pro scores

A Cursor study found that coding agents are retrieving known fixes at runtime rather than actually solving problems, which means SWE-bench Pro scores are being artificially inflated through contamination. This matters because the whole industry uses these numbers to compare agents and justify pricing. If the benchmark is broken, so is the story everyone’s been telling about coding agent progress.

Anthropic Mythos cleared for 100-plus US companies and agencies

The Trump administration has authorised Anthropic’s Mythos 5 for use across more than 100 companies and government agencies, including non-American employees. Paired with the GPT-5.6 access situation, it’s clear the US government now sees frontier models as something to be managed, not just procured. The Anthropic versus OpenAI framing is increasingly less useful when both are operating inside the same political framework.

OpenAI’s Jalapeño chip signals a real shift away from Nvidia

OpenAI’s custom inference chip, built with Broadcom, is a direct response to Nvidia’s margins and the cost of running models at scale. Google, Apple, and SpaceX are all doing the same thing. This isn’t posturing. When your inference costs exceed your headcount, building your own silicon is just arithmetic.

AI startup Lindy drops Claude for Deepseek to survive

Lindy’s CEO says switching from Claude to Deepseek was a survival decision after AI costs overtook personnel costs. This is the quiet story behind all the flagship model launches. For product companies running inference at volume, cost is the feature.

Related