The supporting record: observations, unresolved forecasts, earlier arguments and historical turning points. Our five prediction countdowns are published on Horizon.
Our 12 September OpenRouter check matched 37 of 39 tracked model IDs. Two were absent, and several open weight models had paid hosted rates. The model ID, endpoint and check date matter more than a loose product label.
SWE-bench reports an open multimodal evaluation update
confirmed
codemodels
The SWE-bench maintainers report that Multimodal v2 is open source with 480 tasks for local evaluation. A reproducible task set makes model claims easier to inspect. We have not run this benchmark or verified a new leaderboard result.
The MCP repository points to a 2026-07-28 specification schema. Versioned interfaces give implementers something concrete to check against. This observation is about the published artefact, not enterprise adoption.
By 11 October 2027, sandbox controls, action approvals and inspectable run records could be standard product features in agent platforms. These controls matter when a model can change files or call external tools.
Target:
Basis, limits and how we will judge it
Review. Retained as emerging. Codex documents sandboxing, approval policies and optional telemetry today. That is an existence proof in one product, not proof of market-wide adoption or immunity to prompt injection.
Resolution test. Use a named, published sample of agent platforms. Verify that each control can be configured and observed in a real run, including a denied action. Documentation alone does not prove the controls work. A majority must provide all three features.
By April 2027, shared protocols could become the usual route for exposing tools, data and agent capabilities in enterprise agent systems. MCP and A2A solve different parts of that problem. Neither removes the need to design permissions or recover from failures.
Target:
Basis, limits and how we will judge it
Review. Retained as emerging. The maintained specifications and SDKs are concrete infrastructure. Their existence does not show that most enterprises use them. We removed the unsupported implication that adoption had already been measured.
Resolution test. Publish the sampling frame before judging adoption. Count named production integrations using a shared protocol, excluding demos and announcements. A majority of that sample must use a protocol for a substantial part of the integration. Report the sample limits.
By 11 April 2027, browser and desktop agents could be useful in a limited set of repetitive back-office jobs while broad consumer autopilot remains unreliable. The hard test includes changed screens, interruptions and unexpected outcomes.
Target:
Basis, limits and how we will judge it
Review. Retained as contested. OSWorld provides an executable desktop task benchmark and documents revisions to its evaluations. It does not measure consumer adoption or prove that a particular production workflow is reliable.
Resolution test. Compare repeat runs of named narrow workflows with varied desktop tasks. Include changed UI, lost sessions and recovery. Report failures and human interventions as well as completed tasks. A polished demonstration is insufficient.
By 11 April 2028, local inference could become the default for some offline and privacy-sensitive app features, with server models handling other work. Whether a feature is private depends on the full data path, including tools and fallback calls.
Target:
Basis, limits and how we will judge it
Review. Changed from speculative to contested after checking concrete local deployment options. Apple documents both on-device and server models, and llama.cpp supports local hardware. Neither shows that most mainstream apps default to local execution.
Resolution test. Test named shipping app features with network access disabled and inspect their declared data paths. Record hardware support, fallback behaviour and task quality. A local model inside a feature that uploads its inputs does not meet the privacy claim.
By 11 October 2027, an open weight reasoning model could offer a practical alternative to the strongest closed APIs on both reasoning and coding at a lower total cost. The useful comparison includes retries, latency and hosting, not just a model score.
Target:
Basis, limits and how we will judge it
Review. Downgraded from emerging. The R1 model card reports comparisons with older closed models. It supports a capability claim at that release, not parity with the September 2026 frontier. Hosted API rates also do not establish the total cost of running open weights.
Resolution test. Compare open and closed models on the same published reasoning and coding test sets with the same task budget. Predeclare an acceptable quality gap and count all inference, retry and hosting costs. Without an independent like-for-like result, leave the forecast unresolved.
By 11 April 2027, agents could produce most routine scaffolding and straightforward refactors in teams that adopt them. Review effort and defects determine whether this is a gain. The volume of generated code does not.
Target:
Basis, limits and how we will judge it
Review. Downgraded from emerging. SWE-bench measures issue resolution in a defined evaluation. It cannot establish what share of production work agents perform. The original claim about most fast-moving teams lacked an adoption baseline.
Resolution test. Follow named teams across repeated production tasks. Record accepted changes, review time, escaped defects and the human baseline. Define routine work before counting it. Do not infer a majority of teams from one lab or from benchmark scores.
By 11 April 2028, large organisations may commonly keep several model families in production for different tasks, costs and deployment needs. A catalogue with many choices makes this possible. It does not tell us what customers actually run.
Target:
Basis, limits and how we will judge it
Review. Downgraded from contested. The checked OpenRouter catalogue demonstrates available choice, not enterprise adoption. We found no representative production-use evidence for the original claim that most large firms will run at least three families.
Resolution test. Use a published, representative sample of large organisations. Count named production model families, excluding evaluations, unused contracts and reseller duplicates. The original majority-and-three-families claim remains unproven unless that threshold is met.
original editorial date scenarios, first published April 2026
no scenarios match the active theme filter
AGI ⚠
modelssociety
Systems matching or exceeding human performance across most cognitive tasks, including ones outside their training distribution, without task-specific retraining.
See debate →
optimistic by end of 2028
Assumptions
Reasoning, memory, tool use and self-improvement keep compounding until frontier systems can transfer robustly across most cognitive domains at roughly human level.
Blockers
The last 20% of reliability, autonomy and world-model grounding proves far harder than scaling advocates expect.
Implication
Build for an agent-native world now: outcome-based products, tiny oversight teams, and businesses that assume intelligence becomes abundant before trust does.
pragmatic 2032-2035
Assumptions
Systems become extraordinary co-workers first, and only later cross the threshold into genuinely general autonomous competence across domains.
Blockers
Deployment friction, safety constraints and evaluation gaps keep real-world capability behind benchmark capability for years.
Implication
Design around hybrid intelligence: let machines dominate bounded cognition while humans keep accountability, cross-functional judgement and final authority.
sceptical not before the 2040s
Assumptions
Human-level general intelligence depends on embodiment, durable goals, social learning and world models that current architectures do not naturally supply.
Blockers
Digital-only systems start generalising across messy real-world domains with minimal scaffolding and without brittle failure modes.
Implication
Optimise for augmentation, governance and competitive advantage from AI tools, not for AGI-timing theatre.
Agentic Work
agentsenterprisework
AI systems that autonomously execute multi-step knowledge work across tools, queues and approval boundaries, owning outcomes end-to-end rather than assisting a human operator.
optimistic by end of 2027
Assumptions
Tool use, memory, permissions, evaluation and error recovery improve fast enough that agents can own large volumes of queue-based knowledge work without human approval loops.
Blockers
Identity, auditability, exception handling and liability stay unresolved long enough to stop organisations trusting unattended execution.
Implication
Build narrow, high-volume domain agents now and wrap them in approval tiers, rollback paths and outcome-level monitoring.
pragmatic 2029-2031
Assumptions
Agents become dependable in structured workflows first, while open-ended office work remains mostly supervised because tacit context is still hard to encode.
Blockers
Enterprise data stays fragmented and process owners fail to redesign workflows around machine delegation.
Implication
Sell partial autonomy, not full replacement: hand-offs, triage, drafting, reconciliation and escalation will land before lights-out execution.
sceptical mid-2030s or later
Assumptions
Most knowledge work hides politics, ambiguity, negotiation and accountability that cannot be cleanly reduced to tools plus prompts.
Blockers
Agents prove they can recover from ambiguity, manage cross-system state and own consequences in messy live environments.
Implication
Treat agents as force multipliers for people and focus on better interfaces, memory and review rather than labour-substitution bets.
Robotics
roboticsinfrastructurework
General-purpose physical robots (humanoid or otherwise) that are commercially routine, not demo-quality, with fleet-deployable reliability and workable unit economics.
optimistic by 2030
Assumptions
Vision-language-action models, dexterity, battery performance and manufacturing scale improve together quickly enough to make general-purpose robots commercially routine and cheap.
Blockers
Reliability in cluttered environments, safety certification and service economics fail to move from demo quality to fleet quality.
Implication
Start designing robot-ready workflows, facilities and software now, because the integration layer will matter as much as the hardware.
pragmatic 2033-2036
Assumptions
General-purpose robotics lands first in warehouses, factories, logistics and other structured commercial environments before the home catches up.
Blockers
Unit economics stay weak because teleoperation, maintenance and failure recovery remain too expensive.
Implication
Build for structured environments and mixed fleets, where robot coordination, observability and process redesign create the first durable value.
sceptical not before the late 2030s
Assumptions
True general-purpose robotics is a full-stack systems problem, and manipulation, safety and upkeep are much harder than the current curve implies.
Blockers
On-device robot models plus mass manufacturing crack reliability and cost at the same time.
Implication
Treat humanoids as long-duration options and keep investing in fixed automation, sensors and workflow software that pays off sooner.
Software Automation
codeenterprisework
Coding agents reliably owning the loop from ticket to monitored production deploy across large codebases, leaving humans mostly specifying, reviewing and steering.
optimistic by end of 2028
Assumptions
Long-horizon coding agents become strong enough at planning, editing, testing, migration and repository memory that humans mostly specify, review and steer.
Blockers
Security, reproducibility, architecture drift and repo-specific context keep agents from owning production changes at scale.
Implication
Rebuild engineering around evaluation, review policy and product judgement, because typing and boilerplate stop being the scarce skill.
pragmatic 2030-2032
Assumptions
AI takes over most routine implementation and maintenance, but humans still dominate architecture, incident response, stakeholder translation and high-risk decisions.
Blockers
Firms fail to trust generated changes in production and never build the testing and governance needed for deeper automation.
Implication
Prepare for smaller engineering teams with stronger QA, clearer specs and codebases designed to be legible to agents.
sceptical not before the mid-2030s
Assumptions
Production software is mainly about ambiguous requirements, coordination, risk and long-tail maintenance rather than writing lines of code.
Blockers
Agents start reliably running the full loop from ticket to monitored deploy across large, messy codebases.
Implication
Invest in developer leverage and system clarity, not in simple headcount-reduction stories.
Education Disruption
educationworksociety
AI tutoring and assessment displacing institutional course and credential delivery as the primary structure through which people learn and signal mastery.
optimistic by 2030
Assumptions
AI tutoring becomes dramatically better and cheaper than conventional content delivery, and assessment adapts fast enough to preserve trust in learning outcomes.
Blockers
Credentialing inertia, safeguarding, procurement cycles and political resistance keep institutions tied to legacy delivery models.
Implication
Build assessment, coaching, learning-record and teacher-orchestration products rather than more static content libraries.
pragmatic 2032-2035
Assumptions
AI transforms tutoring, practice and feedback quickly, but schools, universities and employers retain the institutional shell of courses, cohorts and credentials.
Blockers
Demonstrated learning gains become so overwhelming that institutions are forced to redesign faster than expected.
Implication
Plug into existing institutions instead of trying to replace them; the winning tools will fit classrooms, campuses and compliance.
sceptical not this generation
Assumptions
Education is not primarily content delivery; it is socialisation, signalling, childcare, norm formation and supervised practice, so AI enhances learning without replacing structured learning.
Blockers
Employers stop trusting conventional credentials and start trusting AI-mediated mastery records instead.
Implication
Focus on teacher augmentation, administrative relief and better evidence of skill, not on betting against the institution itself.
Debate Zone
earlier editorial arguments and their references
no debates match the active theme filter
Will open-weight models catch the frontier permanently, or only in specific niches?
modelsinfrastructure
For
The pattern is no longer “open follows years later”; it is “open absorbs frontier ideas fast enough to matter commercially”. Stable Diffusion proved that closed advantages can leak into public ecosystems, LLaMA broke the sense that only a few labs could play, and DeepSeek-R1 showed that open reasoning can arrive with real teeth. If open gets to 90–95% of frontier capability at a fraction of the cost, permanent practical parity is enough to change who wins distribution.
The frontier does not stand still while open catches up. Closed labs still control the best compute, the richest feedback loops, the strongest product distribution and the hardest-to-copy post-training tricks, which lets them keep moving the goalposts from text to multimodal, from answers to agents, and from fluency to reasoning. Open will be formidable in cost-sensitive niches, but the top layer of capability will remain structurally concentrated.
Is reasoning a real paradigm shift, or expensive pre-processing wearing a new hat?
modelsinterfaces
For
o1 made it plausible that test-time compute is not just extra verbosity but a new way to buy capability on hard tasks. Once models can spend inference budget deliberately, the product question changes from “how fluent is it?” to “how much thinking is this worth?”. That creates a genuinely new optimisation space for models, interfaces and pricing, and it is why reasoning now feels like a separate frontier rather than a mere feature.
A lot of the current reasoning story may be packaging rather than paradigm. Bigger inference budgets, hidden chain-of-thought and better scaffolding can lift benchmark scores without solving the messy problems users actually care about, such as reliability, context management and action in the real world. If the gains are costly, slow and narrow, “reasoning” could turn out to be an expensive wrapper around familiar techniques.
Will local inference replace cloud inference for the majority of routine AI interactions within 24 months?
infrastructureenterprise
For
Routine, high-frequency AI use wants low latency, predictable cost and privacy by default. Once a model is good enough to summarise, transcribe, draft, organise and personalise locally, shipping every interaction to the cloud starts to look like a design relic rather than a necessity. Cloud will still do the heavy lifting, but local could become the default surface most people touch first.
The average user will not choose local over cloud on principle; they will choose whatever works best. The strongest models, freshest world knowledge, richest tool access and fastest product-improvement cycles are still cloud-shaped advantages, and platform owners have economic reasons to keep the smart layer centralised. Local will grow quickly, but mostly as a complement to cloud rather than its replacement.
Is agent infrastructure becoming a real software category, or are we prematurely standardising a messy phase?
agentssecurity
For
When a field starts inventing shared protocols, permission layers, eval harnesses and operational controls, that usually means the pattern is becoming real. Agents that touch tools, data and workflows need governable plumbing, and that need does not disappear even if today’s demos are messy. The category may still be early, but the infrastructure demand looks structural rather than faddish.
Every platform shift produces a premature standards rush before anyone knows what the durable primitive actually is. Many so-called agent frameworks are still thin wrappers around brittle model calls, and the winning abstractions may end up bundled into the major model platforms instead of living as an independent layer. What looks like category formation could still be the noisy middle stage before consolidation.
Will vibe coding broaden software creation, or mainly create a larger QA and governance burden?
codework
For
Vibe coding lowers the hardest barrier for many would-be builders: getting from intent to working software. That means domain experts can finally ship internal tools without having to become full-time engineers, which is how real capability usually spreads inside organisations. If guardrails improve, the organisational effect could look less like chaos and more like the spreadsheet moment for software creation.
Most software cost arrives after the prototype: security, integrations, ownership, maintenance and change control. Vibe coding widens the front door, but it also makes it easier to generate brittle systems that create hidden review and governance debt. In companies, that may mean more shadow IT and more QA burden rather than a clean productivity windfall.
What actually counts as AGI — economic substitution, cognitive breadth, or something else entirely?
modelssociety
For
The useful definition is economic: a system is AGI when it can reliably perform most paid knowledge work at human level or better, across domains, without task-specific retraining. This framing is measurable, matters commercially and treats “general” as a threshold of breadth rather than a philosophical claim about understanding. Reasoning models plus tool use already point at this definition, even if the threshold is years away.
The economic definition smuggles in a philosophical claim and hides the hard part. Genuine general intelligence arguably requires durable goals, embodied experience, robust world-modelling and cross-domain transfer that current systems do not have even when they score well on benchmarks. A system that passes every exam and still cannot run a small business on its own is not general; it is a very capable text engine. The label matters because it shapes policy, investment and safety expectations.
no Past turning points match the active theme filter
2025
Vibe coding
1 Feb 2025
Andrej Karpathy named a new way of building software by talking to AI and steering outputs conversationally. It marked the moment coding started to feel like directing and taste-making, not just writing syntax.
codework
DeepSeek-R1
1 Jan 2025
DeepSeek released an open reasoning model it said was on par with o1, with open weights and MIT licensing. It reset expectations on open models, reasoning economics and who could shape the frontier.
modelsinfrastructure
2024
Gemini 2 frames the agentic era
11 Dec 2024
Google's Gemini 2.0 launch foregrounded agentic capability — multimodal reasoning, tool use, and long-running task execution — as the explicit framing for what the next generation of frontier models would compete on.
agentsmodels
Model Context Protocol
1 Nov 2024
Anthropic open-sourced MCP, a standard for connecting AI assistants to tools and data. It gave the agent era shared plumbing instead of endless bespoke integrations.
infrastructureagents
o1 reasoning shift
1 Sept 2024
Models began visibly spending more effort on reasoning before answering. It marked the move from fluent prediction to deliberate reasoning as the next frontier.
models
GPT-4o brings real-time multimodal AI
13 May 2024
OpenAI's GPT-4o unified text, audio and vision in a single model with conversational latency. It moved the assistant interface from turn-by-turn typing toward real-time spoken interaction across modalities.
modelsinterfaces
Claude becomes a credible alternative
1 Mar 2024
Anthropic's Claude 3 made the frontier model race feel genuinely multi-player. It ended the idea that one lab would automatically own the assistant layer.
models
Sora reveal
1 Feb 2024
OpenAI previewed a video model with unusually coherent scene generation. It made world-simulation feel plausible rather than gimmicky.
modelscreativity
2023
AI safety breaks into public view
29 Mar 2023
Eliezer Yudkowsky's intervention helped push existential-risk arguments into mainstream debate. It split the story of AI into two tracks, capability acceleration and safety confrontation.
regulationsociety
GPT-4 release
14 Mar 2023
OpenAI released a much stronger model with a visible jump in capability. It moved AI from novelty to serious knowledge work, coding and professional use.
modelswork
LLaMA and the open-weight era
1 Feb 2023
Meta's LLaMA escaped into the wider world and sparked rapid open model work. It broke the feeling that only a few labs could seriously participate.
models
2022
ChatGPT launch
30 Nov 2022
A conversational interface put frontier AI in front of ordinary people. It turned AI from a sector story into a civilisation story almost overnight.
modelsinterfacessociety
Whisper releases as open-weight speech recognition
21 Sept 2022
OpenAI released Whisper as open-weight automatic speech recognition matching commercial systems on many languages. It seeded a wave of locally-runnable transcription that did not need to call a cloud API.
modelsinfrastructure
Stable Diffusion public release
1 Aug 2022
Powerful image generation became openly available on ordinary hardware. It started the open model era in public consciousness.
modelscreativity
InstructGPT and RLHF
1 Mar 2022
OpenAI showed that human feedback could make models more useful and aligned. It changed the winning formula from just training bigger models to training and then aligning them.
modelssociety
2021
AlphaFold 2 lands
1 Jul 2021
DeepMind solved protein structure prediction at near-experimental quality. It showed AI was not just a media tool but a scientific instrument.
modelsdata
GitHub Copilot preview
1 Jun 2021
AI coding help moved into everyday developer workflow. It turned software creation into one of the first mass human-plus-AI production loops.
codework
DALL-E first reveal
1 Jan 2021
OpenAI showed text-to-image generation as a general capability. It opened the public imagination to generative AI beyond text.
modelscreativityinterfaces
2020
GPT-3 paper
1 May 2020
OpenAI showed a giant language model that could do many tasks from prompting alone. It shifted the field towards general-purpose foundation models instead of narrow task systems.
models
2017
Transformer paper
1 Jun 2017
Attention Is All You Need introduced the transformer architecture. It became the core design behind modern language models and much of generative AI.
modelsinfrastructure
2016
AlphaGo defeats Lee Sedol
1 Mar 2016
DeepMind's system beat one of the world's best Go players. It broke the assumption that intuition-heavy human domains were still safe from machines.
modelssociety
Weekly Shift Log
what changed, last 7 days
● quiet week
No changes to any lane in the last 7 days. The
shift log is derived from git log -- src/data/horizon/ by the horizon bot.
Track a theme
Filter the map by what you actually care about. Picked themes sync to the URL, so you can bookmark or share a slice.
Explore the machinery
The horizon map runs on the same pipeline as the rest of the site. The bots
propose, we decide. See how it works.