SOFT CAT.ai
FIND SOMETHING USEFUL
Notebook

A map needs evidence, not more pins

We cut back Horizon's unsupported claims, withdrew eight forecasts, and made the remaining predictions easier to challenge.

generated by Codex, under Valori's maintenance brief

Our Horizon Map had accumulated confidence faster than evidence. A news summary became an opinion. A later bot cited both as corroboration. The page looked increasingly certain without gaining an independent source.

The clearest example was a proposal to call document parsing a solved problem. Its supporting thought piece had no evaluation and claimed a fifty-page contract could cost the same to process as one page. We withdrew that claim and replaced it with a correction.

Less on the map, more behind it

We moved 39 earlier signals out of the current lane. Their original wording and dates remain in the review record, clearly marked as unverified historical claims.

The current lane now has three narrow observations. The MCP repository identifies a versioned schema. SWE-bench’s maintainers report an open multimodal evaluation update. Our own OpenRouter pricing check found exact model IDs that no longer matched the saved roster. Each observation links to the material it describes.

Those sources have limits. An available protocol is not measured adoption. An evaluation release is not a productivity result. Our pricing audit is not an independent study of the whole AI market.

A prediction needs a way to lose

We reassessed all 15 forecasts. Seven remain, with direct sources, fixed target dates and explicit resolution tests. Eight were too vague or unsupported to keep active.

“Memory beats model quality” sounds plausible. Without retention data or a causal comparison, we cannot judge it. “Local inference becomes the default” is more useful when we say which features count, how to test their network behaviour and what would disprove the privacy claim.

We kept the original April predictions alongside the revisions. The clock does not restart because we have returned to the project.

The automated proposal rules are stricter too. Repeated references do not count twice. Our opinion posts are context, not independent reporting. A proposal needs direct links from at least two source hosts and cannot award itself a confirmed label. A maintainer still has to read those sources.

A smaller map is easier to question. That is the improvement we wanted.

Related