#benchmarks
7 posts tagged benchmarks.
News & Updates
AI digest: models that fail, models that return, and one very embarrassing benchmark
This week: AI agents go broke running fake companies, Ford learns AI can't replace experience, and Anthropic's Mythos 5 gets cleared for critical infrastructure.
AI digest: cheating models, restricted flagships, and the race to the edge
GPT-5.6 Sol cheats on tests, Anthropic's flagship slowly returns, and Liquid AI puts a capable model on a Raspberry Pi.
AI digest: government hands on the throttle
GPT-5.6 launches under government-controlled access, benchmark fraud gets exposed, and the custom chip race heats up.
DeepSeek R1 vs Claude 3.5: a head-to-head on real tasks
Ran both models through the same set of coding and reasoning tasks. Results were closer than expected.
Testing Kimi k1.5: the reasoning model nobody's talking about
Moonshot AI's Kimi k1.5 quietly dropped and it's genuinely impressive on long reasoning tasks.
Thoughts
Agentic coding models just made software engineering a spectator sport
When models hit 70% on complex coding benchmarks, we're not optimising development anymore, we're just watching machines do our jobs.
Agent benchmarks are just unit tests for unpredictable systems
We're measuring agent performance like it's deterministic software when the whole point is emergent behaviour.