The green tick was checking too littleOur JSON tool could pass an output while ignoring the constraint that mattered. We replaced the checker, then tested the failure cases.13 Sept 2026field-notesevaluationstructured-outputreliability
A map needs evidence, not more pinsWe cut back Horizon's unsupported claims, withdrew eight forecasts, and made the remaining predictions easier to challenge.12 Sept 2026field-noteshorizonevaluationcorrections
Shorter prompts still need a fair testA shorter system prompt is a candidate to test. One reported reduction cannot establish that careful instructions have stopped mattering.3 Jul 2026promptingmodel-behaviourclaude-codeevaluationcorrections
Document parsing still needs testsCorrection: our original post called document parsing solved without presenting an evaluation. Here is the standard that claim should have met.25 Jun 2026ocrdocument-intelligenceevaluationcorrections
Context size needs a retrieval testFitting a project into a model's input is useful. Finding the right facts and producing a correct answer still need to be measured.14 Feb 2026context-windowevaluationmodelscorrections
Compare answers against a stated rubricJudge specific requirements against a reference without inventing an overall quality score.evaluationmodel-comparisongrounding
Reconcile an AI bill before proposing savingsCalculate the supplied text-token charges and separate measured usage from unpriced optimisation ideas.costtokensevaluation
Shorten a prompt without losing a ruleCompress repeated wording while preserving explicit constraints and showing how to check the revision.promptingeditingevaluation
Summarise a source, keep its caveatsProduce a short summary that preserves denominators, uncertainty and the limits of the supplied evidence.summarisationgroundingevaluation