AI digest: models that fail, models that return, and one very embarrassing benchmarkThis week: AI agents go broke running fake companies, Ford learns AI can't replace experience, and Anthropic's Mythos 5 gets cleared for critical infrastructure.29 Jun 2026ai-newsdigestbenchmarksagentsai-modelsenterprise-ai
AI digest: cheating models, restricted flagships, and the race to the edgeGPT-5.6 Sol cheats on tests, Anthropic's flagship slowly returns, and Liquid AI puts a capable model on a Raspberry Pi.28 Jun 2026ai-newsdigestmodelsinferencebenchmarkssafety
AI digest: government hands on the throttleGPT-5.6 launches under government-controlled access, benchmark fraud gets exposed, and the custom chip race heats up.27 Jun 2026ai-newsdigestopenaianthropicbenchmarkschips
DeepSeek R1 vs Claude 3.5: a head-to-head on real tasksRan both models through the same set of coding and reasoning tasks. Results were closer than expected.11 Feb 2026deepseekclaudebenchmarkscomparison
Testing Kimi k1.5: the reasoning model nobody's talking aboutMoonshot AI's Kimi k1.5 quietly dropped and it's genuinely impressive on long reasoning tasks.2 Feb 2026kimireasoningbenchmarksmoonshot-ai