A map needs evidence, not more pinsWe cut back Horizon's unsupported claims, withdrew eight forecasts, and made the remaining predictions easier to challenge.12 Sept 2026field-noteshorizonevaluationcorrections
Shorter prompts still need a fair testA shorter system prompt is a candidate to test. One reported reduction cannot establish that careful instructions have stopped mattering.3 Jul 2026promptingmodel-behaviourclaude-codeevaluationcorrections
Measure cost changes before assigning motivesA changed bill needs a traceable explanation. Our earlier essay made allegations without providing the records needed to assess them.2 Jul 2026ai-trustmodel-pricingobservabilitycorrections
Document parsing still needs testsCorrection: our original post called document parsing solved without presenting an evaluation. Here is the standard that claim should have met.25 Jun 2026ocrdocument-intelligenceevaluationcorrections
Hybrid routing should show where a request goesIf a tool can switch between a local model and a hosted one, the interface should make that choice understandable.6 Jun 2026hybrid-inferencerouting-logicprivacyinterfacescorrections
Agent controls need tests at the boundaryWe can test what an agent is allowed to touch without pretending to know everything about why it acts.19 Mar 2026agent-securityruntime-safetypermissionsreliabilitycorrections
Context size needs a retrieval testFitting a project into a model's input is useful. Finding the right facts and producing a correct answer still need to be measured.14 Feb 2026context-windowevaluationmodelscorrections