THE LONGER RECORD
Thoughts & corrections.
Earlier writing that still has a useful point, revised where its claims went beyond the evidence. Corrections explain what changed and link to the earlier text.
For documented experiments and maintenance, start with the notebook.
CHANGING OUR MINDS IN PUBLIC
The correction record
Shorter prompts still need a fair test
Removed an unsupported 80 per cent prompt-reduction claim and the conclusion that prompting had become obsolete.
Originally published 3 Jul 2026Measure cost changes before assigning motives
Withdrew unsupported claims about hidden monitoring, a 40 per cent token increase and deliberate price manipulation.
Originally published 2 Jul 2026Hybrid routing should show where a request goes
Removed the mental-health metaphor and unsupported claims about how all hybrid systems behave. The useful issue is request routing and disclosure.
Originally published 6 Jun 2026Agent controls need tests at the boundary
Withdrew blanket claims that security frameworks are useless and that a confused agent will inevitably escape a sandbox.
Originally published 19 Mar 2026The original SOFT CAT pipeline
Reframed the launch article as project history. Removed claims that all six bots ran this morning, corrected the treatment of schedules and linked the current records. The separate server is not verified by a new website build.
Originally published 13 Mar 2026Context size needs a retrieval test
Removed an undocumented whole-project test and the blanket conclusion that a larger context window beats a more capable model.
Originally published 14 Feb 2026Document parsing still needs tests
Withdrew the solved-problem claim and the unsupported claim of equal cost and latency across document lengths.
Originally published 25 Jun 2026