SOFT CAT.ai
FIND SOMETHING USEFUL
← Back to Horizon

Editorial review ·

A smaller map.
A clearer trail.

A forecast needs a deadline and a way to be wrong. A source needs to support the claim we attach to it.

7active forecasts
8withdrawn forecasts
39earlier signals archived

What changed

The old current lane promoted summaries and opinions into broad claims about adoption. Some may prove useful, but their source chains did not support that level of confidence. We moved them out of the current map and kept the historical record below.

The current lane now contains narrow observations linked to their source material. A document can establish that a specification exists. A benchmark release can establish that an evaluation is available. Neither proves market share, productivity or safety.

We reassessed all 15 forecasts. Seven now have direct sources, dated targets and explicit resolution tests. Eight were too vague or unsupported to keep active. Withdrawal is an editorial decision, not a claim that the opposite outcome is true.

Confidence labels are qualitative. The three worldview scenarios elsewhere on Horizon are editorial hypotheses, not a survey or expert consensus. We keep unresolved questions unresolved.

Inspect the original forecast snapshot ↗
Open-weight reasoning reaches practical frontier parityemerging → contested

Original claim · 11 April 2026

Within 18 months, at least one open-weight reasoning model will sit within striking distance of the best closed APIs on public reasoning and coding evals at materially lower cost. The DeepSeek-R1 to Gemma 4 arc suggests the open curve is steepening faster than frontier labs can preserve a clean capability moat.

Review · 12 September 2026

Downgraded from emerging. The R1 model card reports comparisons with older closed models. It supports a capability claim at that release, not parity with the September 2026 frontier. Hosted API rates also do not establish the total cost of running open weights.

Current claim, sources and resolution test ↗
Protocols beat bespoke agent plumbingemerging → emerging

Original claim · 11 April 2026

By April 2027, most serious enterprise agent stacks will expose tools and data through an MCP/A2A-style protocol layer rather than mostly bespoke connectors. Once agents need many tools, shared plumbing becomes cheaper and more governable than custom glue.

Review · 12 September 2026

Retained as emerging. The maintained specifications and SDKs are concrete infrastructure. Their existence does not show that most enterprises use them. We removed the unsupported implication that adoption had already been measured.

Current claim, sources and resolution test ↗
Agent security becomes a platform layeremerging → emerging

Original claim · 11 April 2026

Within 18 months, sandboxing, approval gates and audit logs will be standard features in serious agent platforms. Tool-connected agents turn prompt injection from a content problem into an operational-security problem.

Review · 12 September 2026

Retained as emerging. Codex documents sandboxing, approval policies and optional telemetry today. That is an existence proof in one product, not proof of market-wide adoption or immunity to prompt injection.

Current claim, sources and resolution test ↗
Coding agents own the easy middleemerging → speculative

Original claim · 11 April 2026

Within 12 months, AI agents will produce most greenfield boilerplate, test scaffolding and straightforward refactors in fast-moving software teams. The capability is already good enough for bounded software work, while the hard part remains review, architecture and production judgement.

Review · 12 September 2026

Downgraded from emerging. SWE-bench measures issue resolution in a defined evaluation. It cannot establish what share of production work agents perform. The original claim about most fast-moving teams lacked an adoption baseline.

Current claim, sources and resolution test ↗
Computer-use agents find narrow fitcontested → contested

Original claim · 11 April 2026

Within 12 months, browser and desktop agents will stick in a small set of repetitive back-office workflows but remain too brittle for broad consumer autopilot. The capability curve is real, but permissions, exception handling and reliability are still improving slower than the demos.

Review · 12 September 2026

Retained as contested. OSWorld provides an executable desktop task benchmark and documents revisions to its evaluations. It does not measure consumer adoption or prove that a particular production workflow is reliable.

Current claim, sources and resolution test ↗
Enterprise AI stays multi-modelcontested → speculative

Original claim · 11 April 2026

Within 24 months, most large enterprises will run at least three model families in production rather than standardise on one provider. Bedrock, Vertex and Azure are all training buyers to purchase optionality instead of allegiance.

Review · 12 September 2026

Downgraded from contested. The checked OpenRouter catalogue demonstrates available choice, not enterprise adoption. We found no representative production-use evidence for the original claim that most large firms will run at least three families.

Current claim, sources and resolution test ↗
Local inference becomes the private defaultspeculative → contested

Original claim · 11 April 2026

Within 24 months, many privacy-sensitive or offline AI features in mainstream apps will default to local inference, with cloud models reserved for heavier reasoning. The model-size curve is falling fast enough that latency, privacy and cost can outweigh raw frontier quality.

Review · 12 September 2026

Changed from speculative to contested after checking concrete local deployment options. Apple documents both on-device and server models, and llama.cpp supports local hardware. Neither shows that most mainstream apps default to local execution.

Current claim, sources and resolution test ↗
Search becomes an assistant surfaceWithdrawn · 2026-09-12

Withdrawn because power users, query categories and the adoption denominator were undefined. Product launches cannot resolve a claim about where most searches begin.

Original claim · 2026-04-11

By late 2027, planning, shopping and research queries will more often begin in assistant-style surfaces than in a blank search box for power users. Search is already absorbing reasoning, follow-up dialogue, personal context and agentic steps.
Memory beats model deltaWithdrawn · 2026-09-12

Withdrawn because no retention data or causal comparison supported the claim that memory matters more than model quality. A plausible product opinion is not a measurable forecast.

Original claim · 2026-04-11

Within 12 months, assistant retention will depend more on memory, projects and connectors than on small differences in base-model quality. Once models are all good enough, continuity starts to matter more than eloquence.
Robotics pays off at work before homeWithdrawn · 2026-09-12

Withdrawn because clearer economic value had no defined cost basis, sample or comparison between industrial and domestic use. It needs a narrower outcome and deployment evidence.

Original claim · 2026-04-11

By 2028, AI-driven robots will create clearer economic value in warehouses, factories and structured commercial settings than in ordinary homes. Manipulation is improving quickly, but the home is still the messiest possible deployment environment.
Vibe coding goes corporateWithdrawn · 2026-09-12

Withdrawn because useful internal tools and organisational uptake were not defined. It also overlapped the coding forecast without adding a separate test.

Original claim · 2026-04-11

Within 18 months, non-engineers in medium and large firms will ship useful internal tools through conversational builders and agentic IDEs. The interface to software is getting easier faster than organisations can redesign review, integration and security.
Education adapts assessment before credentialingWithdrawn · 2026-09-12

Withdrawn because schools, employers and credentialing were combined into one claim without a common measure. This needs institution-specific evidence.

Original claim · 2026-04-11

Within 24 months, schools and employers will rely more on supervised, oral and process-based assessment because AI tutoring is easier to absorb than AI credentialing. Teaching is easier to displace than trusted signalling.
AI video becomes normal pre-productionWithdrawn · 2026-09-12

Withdrawn because routine pre-production and a clean substitute for premium live action had no thresholds or comparable sample. Tool capabilities alone cannot establish production practice.

Original claim · 2026-04-11

Within 18 months, AI video will be routine for pitches, previs, explainers and ad variants but not a clean substitute for premium live-action production. The tools are improving fast, but taste, control, rights and production workflows still matter.
Labelling outruns licensingWithdrawn · 2026-09-12

Withdrawn because the forecast compared jurisdictions and platforms using undefined meanings of enforce and settle. Article 50 of the EU AI Act provides a concrete transparency rule, but cannot resolve that cross-market comparison.

Original claim · 2026-04-11

Within 24 months, more jurisdictions and platforms will enforce synthetic-media labelling or provenance rules than will settle training-data licensing rules. Transparency is easier to operationalise than copyright economics.
Regulation stays regionally forkedWithdrawn · 2026-09-12

Withdrawn because rule-heavy, lighter-touch and no clean global settlement were too broad to resolve consistently. Specific legal duties belong in a dated jurisdiction-by-jurisdiction review.

Original claim · 2026-04-11

By 2028, AI regulation will still be regionally split, with the EU more rule-heavy, the UK lighter-touch, and no clean global settlement on copyright or frontier-model duties. The compliance machinery is diverging faster than international consensus is forming.

These records retain their original wording and dates. Their old confidence labels do not represent the outcome of this review. They are excluded from the current lane.

Open-weight reasoning models match the frontier and undercut the price2026-04-10 · archived

GLM-5.1 (Z.AI) ships as a 754B open-weight agentic model that hits SOTA on SWE-Bench Pro and sustains 8-hour autonomous execution at roughly a third of frontier API cost. Three distinct moats collapse into one release: open weights, frontier capability, agentic competence. Direct continuation of the DeepSeek-R1 thread from January.

Original label: confirmed. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
The agent production-ops layer is crystallising2026-04-10 · archived

Three independent launches in three days, each targeting a different layer of the agent production stack: BotCTL (process management, billed as systemd for agents), OnCell.ai (per-user isolation), Relvy (automated on-call runbooks, YC F24). These are problems that only matter once agents are actually being deployed in production, which means agents are actually being deployed in production.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Tool-calling and agent-messaging protocols are hardening into infrastructure2026-04-10 · archived

AgentDM (agent-to-agent over MCP and A2A), QVeris (10k capabilities discoverable via one protocol), Postagent (Postman-style CLI for agents), and ZeroID (OIDF-based agent identity) all landed within two days. The pattern is the same shape MCP started in late 2024. Agent infrastructure stops being framework wars and starts being shared plumbing.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
OS-as-environment is becoming the standard for agent training2026-04-10 · archived

OSGym ships infrastructure to manage 1,000+ OS replicas at $0.23 per day for computer-use agent research. That is not a product launch, it is a capability. Parallel rollouts on real operating systems become economically viable. Astropad Workbench and TUI-use round out the pattern: agents need bodies, and the bodies are computers.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Local-first inference moves from edge case to default for new launches2026-04-10 · archived

Four launches in five days defaulting to local-first inference rather than cloud: Google's offline AI dictation app on iOS using Gemma, Imbue's Bouncer (on-device LLM for Twitter feed control), QVAC SDK (universal JS for local AI), and Meta's EUPE (sub-100M-parameter vision encoder family). Notable that Google itself is shipping offline-first using its own models. That is the shift, not any single launch.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Multimodal reasoning becomes a frontier-race table-stakes capability2026-04-10 · archived

Meta Superintelligence Lab released Muse Spark, a multimodal reasoning model with thought compression and parallel agents. Frontend-VisualQA (Yutori AI) gave coding agents visual verification of their own UI work. EUPE shows compact vision encoders can rival specialists. Multimodal reasoning is no longer a frontier-lab talking point. It is becoming a baseline capability across the stack.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Debug logs and error traces are becoming premium training data sources2026-04-24 · archived

Failed tests and error logs contain information about edge cases and real-world complexity that clean datasets miss. This data is suddenly more valuable than the working code.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI agent interactions turn market dynamics into debugging problems2026-05-01 · archived

When AI agents trade and interact with each other at scale, economic inefficiencies become software bugs we can identify and patch. Markets become programmable systems rather than natural forces.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Coding models cross the 70% benchmark threshold into production territory2026-05-02 · archived

When AI models hit 70% on complex coding benchmarks like SWE-bench, we're crossing from helpful assistants to primary developers. The gap between AI capability and human necessity in software engineering just collapsed.

Original label: confirmed. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI valuations enter sovereign wealth territory as infrastructure becomes geopolitical2026-05-02 · archived

Anthropic heading for a £900B valuation whilst big tech drops £725B on AI infrastructure signals AI companies reaching nation-state scale economics. This isn't venture capital anymore, it's geopolitical infrastructure investment.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Remote AI agents shift development from local to cloud-first2026-05-03 · archived

Cloud-based AI agents are making local development environments obsolete. Remote execution is becoming the default for agent deployment.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Privacy-preserving AI tools move from research to production2026-05-03 · archived

Major AI companies are shipping open-source privacy tools as protection becomes essential for deployment. Privacy is shifting from afterthought to core infrastructure.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Time-aware model architectures become the new baseline2026-05-04 · archived

Models that process time-based data natively aren't just better at audio and video. They represent a fundamental shift in how AI understands sequential information.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI company valuations detach from engineering fundamentals2026-05-04 · archived

When AI startups get billion-dollar valuations before solving basic engineering problems, investment becomes speculation rather than technology assessment.

Original label: confirmed. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Model distillation becomes standard practice for deployment2026-05-04 · archived

Teaching smaller models to mimic larger ones is becoming the default path from research to production. This isn't just optimisation, it's how AI knowledge transfers.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Sparse autoencoders move from research curiosity to production interpretability standard2026-05-05 · archived

Model interpretability is shifting from academic exercise to operational necessity. Companies need to understand what their AI systems are doing, not just whether they work.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Clean training data transforms from commodity to premium service category2026-05-05 · archived

The AI training pipeline is splitting into haves and have-nots based on data quality. Clean data is becoming the new competitive moat.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Voice AI crosses the real-time interaction barrier2026-05-06 · archived

Voice AI is moving from stilted turn-taking to natural conversation flow. This shifts AI from a tool you operate to a participant you converse with.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI notifications become the new engagement layer2026-05-06 · archived

Push notifications for AI jobs turn models from batch processors into always-on services. This transforms AI from a tool you visit to a service that reaches you.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Specification-driven development becomes the production standard for AI agents2026-05-09 · archived

AI coding agents are shifting from prototyping tools to production engineers by requiring formal specifications before execution. This separates serious deployment from experimentation.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Custom networking protocols become competitive moats in AI infrastructure2026-05-09 · archived

AI companies are building proprietary networking protocols to control large-scale training and inference infrastructure. This creates new lock-in mechanisms beyond model weights.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Model packaging shifts from size variants to deployment configurations2026-05-10 · archived

The industry is moving from shipping different model sizes to shipping different deployment configurations of the same model. This changes how we think about model distribution and optimisation.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Real-time interaction becomes the new baseline for AI interfaces2026-05-10 · archived

AI interfaces are crossing from request-response patterns to continuous real-time interaction. This fundamentally changes how we design AI experiences.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Command-line interfaces become the preferred agent interaction pattern2026-06-10 · archived

Chat interfaces are proving inadequate for complex AI agent workflows. Terminal-style interactions force precision and enable proper operational control.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI model repositories face systematic security contamination2026-06-10 · archived

Model repositories are becoming attack vectors with no established cleanup protocols. The infrastructure for safe model distribution is broken.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Real-time streaming responses become competitive battleground2026-06-10 · archived

Every AI company is racing to stream responses faster, creating pressure for instant interaction even when it degrades user experience.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Dual-tier safety models become product differentiation strategy2026-06-11 · archived

Companies are shipping identical models with different safety configurations to serve both enterprise and consumer markets simultaneously. This creates artificial product tiers where safety becomes a feature you pay to customise.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Command-line agents force operational complexity onto developers2026-06-11 · archived

AI agents are gravitating towards command-line interfaces because they force clarity, but this dumps decades of system administration complexity directly onto developers. The terminal is becoming the preferred agent interaction layer.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Hybrid local-cloud routing creates schizophrenic AI systems2026-06-11 · archived

Smart task routing between local and cloud models is creating systems that can't decide what they are. Local inference frameworks are making cloud APIs look obsolete whilst hybrid systems try to bridge both worlds.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Agent self-improvement through memory becomes a live safety concern2026-06-21 · archived

Agents that update their own memory from experience can quietly encode bad behaviour with no checkpoint. The gap between capability and oversight is opening faster than tooling to close it.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Ephemeral, disposable infrastructure becomes the default model for agent identity2026-06-21 · archived

Temporary accounts and blank-slate modes for agents signal a shift away from persistent identity toward throwaway, task-scoped instances. This changes how trust, auth, and billing get handled at the infrastructure layer.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Sub-4B models take on specialist roles previously reserved for large models2026-06-21 · archived

A 3B reasoning model and a retrieval-focused small model appearing in the same radar cycle shows the compression trend moving from general chat into harder tasks. Small is no longer a concession, it is a design choice.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Cloud platforms patch agent infrastructure gaps in real time2026-06-22 · archived

AWS and Cloudflare both shipped agent-focused infrastructure fixes in the same week. Platforms are no longer waiting for standards to settle before solving agent operational problems themselves.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Government AI bans accelerate the adoption they intend to stop2026-06-22 · archived

Banning AI models hands labs free publicity and drives curiosity-driven uptake. The Anthropic ban saga shows regulation can backfire badly when the product is software.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Automated commit governance becomes a distinct product category2026-06-22 · archived

Tools that gate, review, or audit AI-generated commits are shipping as standalone products. As AI writes more code, the enforcement layer around that code is becoming its own market.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI organisational memory quietly deepens vendor lock-in2026-06-24 · archived

As agents learn internal workflows and accumulate context, switching costs compound silently. What looks like a productivity gain becomes structural dependency on a single vendor's memory layer.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
AI coding tools expand vertically into platform and mobile territory2026-06-24 · archived

Cursor adding Git platform features and a mobile app signals that AI coding tools are no longer competing on model quality alone. The battleground is shifting to full developer workflow ownership.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Continuous agent loops shift from product feature to operational liability2026-06-23 · archived

Agents running autonomously in the background are exposing the same failure modes as distributed systems: race conditions, unbounded retries, and no clear ownership. The gap between demo and production is turning into an engineering crisis.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗
Offensive AI threats move from theoretical to imminent on intelligence agency timelines2026-06-23 · archived

Five Eyes agencies are now saying offensive AI threats are months away, not years. That shifts the conversation from policy speculation to active defence preparation.

Original label: emerging. This is a preserved historical claim, not a newly verified fact.

Original references and metadata ↗

What earns a place next

A specific observation, a direct source, a date that means something, and a short explanation of the source's limits. Broader patterns need independent corroboration. Our own summary plus our own opinion is one source chain.

New bot proposals remain proposals. The bot cannot award a confirmed label, count duplicated references as corroboration, or use an opinion post as independent reporting.

Follow the work in the notebook ↗