{"schema":1,"scenarios":[{"id":"scenario-agi","topic":"AGI","themes":["models","society"],"definition":"Systems matching or exceeding human performance across most cognitive tasks, including ones outside their training distribution, without task-specific retraining.","contested":true,"debate_ref":"debate-what-counts-as-agi","optimistic":{"timeframe":"by end of 2028","year":2028,"band":"near","assumptions":"Reasoning, memory, tool use and self-improvement keep compounding until frontier systems can transfer robustly across most cognitive domains at roughly human level.","blockers":"The last 20% of reliability, autonomy and world-model grounding proves far harder than scaling advocates expect.","implication":"Build for an agent-native world now: outcome-based products, tiny oversight teams, and businesses that assume intelligence becomes abundant before trust does."},"pragmatic":{"timeframe":"2032-2035","year":2033,"band":"mid","assumptions":"Systems become extraordinary co-workers first, and only later cross the threshold into genuinely general autonomous competence across domains.","blockers":"Deployment friction, safety constraints and evaluation gaps keep real-world capability behind benchmark capability for years.","implication":"Design around hybrid intelligence: let machines dominate bounded cognition while humans keep accountability, cross-functional judgement and final authority."},"sceptical":{"timeframe":"not before the 2040s","year":2045,"band":"far","assumptions":"Human-level general intelligence depends on embodiment, durable goals, social learning and world models that current architectures do not naturally supply.","blockers":"Digital-only systems start generalising across messy real-world domains with minimal scaffolding and without brittle failure modes.","implication":"Optimise for augmentation, governance and competitive advantage from AI tools, not for AGI-timing theatre."}},{"id":"scenario-agentic-work","topic":"Agentic Work","themes":["agents","enterprise","work"],"definition":"AI systems that autonomously execute multi-step knowledge work across tools, queues and approval boundaries, owning outcomes end-to-end rather than assisting a human operator.","optimistic":{"timeframe":"by end of 2027","year":2027,"band":"near","assumptions":"Tool use, memory, permissions, evaluation and error recovery improve fast enough that agents can own large volumes of queue-based knowledge work without human approval loops.","blockers":"Identity, auditability, exception handling and liability stay unresolved long enough to stop organisations trusting unattended execution.","implication":"Build narrow, high-volume domain agents now and wrap them in approval tiers, rollback paths and outcome-level monitoring."},"pragmatic":{"timeframe":"2029-2031","year":2030,"band":"near","assumptions":"Agents become dependable in structured workflows first, while open-ended office work remains mostly supervised because tacit context is still hard to encode.","blockers":"Enterprise data stays fragmented and process owners fail to redesign workflows around machine delegation.","implication":"Sell partial autonomy, not full replacement: hand-offs, triage, drafting, reconciliation and escalation will land before lights-out execution."},"sceptical":{"timeframe":"mid-2030s or later","year":2036,"band":"far","assumptions":"Most knowledge work hides politics, ambiguity, negotiation and accountability that cannot be cleanly reduced to tools plus prompts.","blockers":"Agents prove they can recover from ambiguity, manage cross-system state and own consequences in messy live environments.","implication":"Treat agents as force multipliers for people and focus on better interfaces, memory and review rather than labour-substitution bets."}},{"id":"scenario-robotics","topic":"Robotics","themes":["robotics","infrastructure","work"],"definition":"General-purpose physical robots (humanoid or otherwise) that are commercially routine, not demo-quality, with fleet-deployable reliability and workable unit economics.","optimistic":{"timeframe":"by 2030","year":2030,"band":"near","assumptions":"Vision-language-action models, dexterity, battery performance and manufacturing scale improve together quickly enough to make general-purpose robots commercially routine and cheap.","blockers":"Reliability in cluttered environments, safety certification and service economics fail to move from demo quality to fleet quality.","implication":"Start designing robot-ready workflows, facilities and software now, because the integration layer will matter as much as the hardware."},"pragmatic":{"timeframe":"2033-2036","year":2035,"band":"mid","assumptions":"General-purpose robotics lands first in warehouses, factories, logistics and other structured commercial environments before the home catches up.","blockers":"Unit economics stay weak because teleoperation, maintenance and failure recovery remain too expensive.","implication":"Build for structured environments and mixed fleets, where robot coordination, observability and process redesign create the first durable value."},"sceptical":{"timeframe":"not before the late 2030s","year":2038,"band":"far","assumptions":"True general-purpose robotics is a full-stack systems problem, and manipulation, safety and upkeep are much harder than the current curve implies.","blockers":"On-device robot models plus mass manufacturing crack reliability and cost at the same time.","implication":"Treat humanoids as long-duration options and keep investing in fixed automation, sensors and workflow software that pays off sooner."}},{"id":"scenario-software-automation","topic":"Software Automation","themes":["code","enterprise","work"],"definition":"Coding agents reliably owning the loop from ticket to monitored production deploy across large codebases, leaving humans mostly specifying, reviewing and steering.","optimistic":{"timeframe":"by end of 2028","year":2028,"band":"near","assumptions":"Long-horizon coding agents become strong enough at planning, editing, testing, migration and repository memory that humans mostly specify, review and steer.","blockers":"Security, reproducibility, architecture drift and repo-specific context keep agents from owning production changes at scale.","implication":"Rebuild engineering around evaluation, review policy and product judgement, because typing and boilerplate stop being the scarce skill."},"pragmatic":{"timeframe":"2030-2032","year":2031,"band":"mid","assumptions":"AI takes over most routine implementation and maintenance, but humans still dominate architecture, incident response, stakeholder translation and high-risk decisions.","blockers":"Firms fail to trust generated changes in production and never build the testing and governance needed for deeper automation.","implication":"Prepare for smaller engineering teams with stronger QA, clearer specs and codebases designed to be legible to agents."},"sceptical":{"timeframe":"not before the mid-2030s","year":2036,"band":"far","assumptions":"Production software is mainly about ambiguous requirements, coordination, risk and long-tail maintenance rather than writing lines of code.","blockers":"Agents start reliably running the full loop from ticket to monitored deploy across large, messy codebases.","implication":"Invest in developer leverage and system clarity, not in simple headcount-reduction stories."}},{"id":"scenario-education-disruption","topic":"Education Disruption","themes":["education","work","society"],"definition":"AI tutoring and assessment displacing institutional course and credential delivery as the primary structure through which people learn and signal mastery.","optimistic":{"timeframe":"by 2030","year":2030,"band":"near","assumptions":"AI tutoring becomes dramatically better and cheaper than conventional content delivery, and assessment adapts fast enough to preserve trust in learning outcomes.","blockers":"Credentialing inertia, safeguarding, procurement cycles and political resistance keep institutions tied to legacy delivery models.","implication":"Build assessment, coaching, learning-record and teacher-orchestration products rather than more static content libraries."},"pragmatic":{"timeframe":"2032-2035","year":2034,"band":"mid","assumptions":"AI transforms tutoring, practice and feedback quickly, but schools, universities and employers retain the institutional shell of courses, cohorts and credentials.","blockers":"Demonstrated learning gains become so overwhelming that institutions are forced to redesign faster than expected.","implication":"Plug into existing institutions instead of trying to replace them; the winning tools will fit classrooms, campuses and compliance."},"sceptical":{"timeframe":"not this generation","band":"indefinite","assumptions":"Education is not primarily content delivery; it is socialisation, signalling, childcare, norm formation and supervised practice, so AI enhances learning without replacing structured learning.","blockers":"Employers stop trusting conventional credentials and start trusting AI-mediated mastery records instead.","implication":"Focus on teacher augmentation, administrative relief and better evidence of skill, not on betting against the institution itself."}}],"review":{"reviewed_at":"2026-09-13","dates_origin":"2026-04-22","decision":"We reassessed the five futures against the sources below. They support specific capabilities and expose important gaps, but do not calibrate arrival years. We retain the April windows as illustrative scenarios, not measured forecasts or expert consensus. No arrival date has been moved in this review.","futures":[{"id":"scenario-agi","title":"General intelligence","subtitle":"Beyond narrow tasks","question":"Is AGI already here?","timeframes":{"optimistic":"by end of 2028","pragmatic":"2032-2035","sceptical":"not before the 2040s"},"assessment":"The threshold depends on the definition. We retain these windows as prompts for discussion. Neither a definition nor a benchmark result establishes a universal arrival year.","watch":"Independent tests across unfamiliar domains, with the task selection, human comparison, cost and required assistance published.","evidence":[{"title":"Google DeepMind: Levels of AGI","url":"https://deepmind.google/research/publications/66938/","date_label":"ICML 2024","finding":"This framework distinguishes breadth and depth of capability, with autonomy considered separately.","limit":"A framework for describing progress is not a certification that a current model has passed every threshold."},{"title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3","date_label":"2026 benchmark","finding":"Interactive environments test how agents explore, discover goals and adapt to unfamiliar tasks.","limit":"This measures a particular form of adaptation. A score alone cannot establish competence across all cognitive or economic work."}]},{"id":"scenario-agentic-work","title":"Agents at work","subtitle":"Owning whole workflows","question":"When does assistance become ownership?","timeframes":{"optimistic":"by end of 2027","pragmatic":"2029-2031","sceptical":"mid-2030s or later"},"assessment":"Task completion is measurable, but reliable ownership of a workplace process is a broader claim. The dates remain illustrative, with no new adoption probability assigned.","watch":"Repeated production workflows that report completion, human intervention, recovery, cost and failures, including changed tools and missing context.","evidence":[{"title":"METR: task-completion time horizons","url":"https://metr.org/time-horizons/","date_label":"Page updated 8 May 2026","finding":"METR estimates task difficulty at specified success rates, using human completion time as the reference.","limit":"These are mainly well-specified software tasks. A time horizon is not continuous autonomous runtime or proof that a whole job can be automated."},{"title":"METR: metrics of agent ability","url":"https://metr.org/notes/2026-07-24-metrics-of-model-ability/","date_label":"24 July 2026","finding":"The research note compares performance measures that account for time, cost and the human baseline.","limit":"It analyses measurement choices. It does not establish widespread enterprise adoption."}]},{"id":"scenario-robotics","title":"Robotics","subtitle":"Beyond the demonstration","question":"When do robots become routine?","timeframes":{"optimistic":"by 2030","pragmatic":"2033-2036","sceptical":"not before the late 2030s"},"assessment":"The July model cards provide concrete evidence of robotics work and its limits. They do not establish fleet economics or commercial general-purpose reliability. The date windows remain illustrative.","watch":"Independently reported deployments with uptime, task variation, intervention rates, maintenance costs and safety performance over time.","evidence":[{"title":"Google DeepMind: Gemini Robotics ER 2","url":"https://deepmind.google/models/model-cards/gemini-robotics-er-2/","date_label":"30 July 2026","finding":"The model card describes spatial and temporal reasoning, tool orchestration and success detection for physical agents.","limit":"The publisher restricts safety-critical uses. Reasoning capability alone does not demonstrate reliable physical deployment."},{"title":"Google DeepMind: Robotics On-Device 2","url":"https://deepmind.google/models/model-cards/gemini-robotics-on-device-2/","date_label":"30 July 2026","finding":"The card describes local robot action generation and evaluations of manipulation across several robot platforms.","limit":"Access is limited to trusted testers. The card identifies limits on unfamiliar tasks and complex robot control."}]},{"id":"scenario-software-automation","title":"Software creation","subtitle":"From ticket to production","question":"Who owns the finished software?","timeframes":{"optimistic":"by end of 2028","pragmatic":"2030-2032","sceptical":"not before the mid-2030s"},"assessment":"The evidence supports testing useful delegation, with careful measurement of the whole task. It does not establish an arrival year for autonomous production delivery. The original windows remain illustrative.","watch":"Accepted production changes, review time, escaped defects, incident recovery and total cost against a stated human baseline.","evidence":[{"title":"METR: revising its productivity experiment","url":"https://metr.org/blog/2026-02-24-uplift-update/","date_label":"24 February 2026","finding":"METR reports that selection effects make its later developer experiment an unreliable estimate of the current productivity effect.","limit":"This is a reason to improve measurement. It establishes neither universal slowdown nor a reliable universal speedup."},{"title":"METR: task-completion methodology","url":"https://metr.org/time-horizons/","date_label":"Page updated 8 May 2026","finding":"The published methodology makes the task set, human reference and success thresholds inspectable.","limit":"Successful benchmark tasks do not include every responsibility involved in maintaining a live software service."}]},{"id":"scenario-education-disruption","title":"Education","subtitle":"Learning and credentials","question":"Does better tutoring change the institution?","timeframes":{"optimistic":"by 2030","pragmatic":"2032-2035","sceptical":"not this generation"},"assessment":"Better tutoring and replacement of institutional education are different claims. The pragmatic case describes improvement within institutions, so it is a partial transformation rather than fulfilment of the full replacement definition. The dates remain illustrative.","watch":"Independent replication, retained learning, transfer to new problems, effects on teachers, and changes in how employers recognise learning.","evidence":[{"title":"Kestin and colleagues: AI tutoring trial","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12179260/","date_label":"Scientific Reports, 2025","finding":"A randomised trial in undergraduate physics reported higher immediate learning gains with a specifically designed AI tutor than with the comparison lessons.","limit":"A study of selected lessons does not establish long-term learning across subjects, replacement of teachers or a change in credentials."}]}]},"history":[{"id":"baseline-2026-04-22","date":"2026-04-22","kind":"baseline","source":"https://github.com/valorifutures/softcat.ai/blob/3066df13e0493ef1446c9a3e48f307aa8bf30e5f/src/data/horizon/scenarios.json","records":[{"id":"scenario-agi","definition":"Systems matching or exceeding human performance across most cognitive tasks, including ones outside their training distribution, without task-specific retraining.","timeframes":{"optimistic":"by end of 2028","pragmatic":"2032-2035","sceptical":"not before the 2040s"}},{"id":"scenario-agentic-work","definition":"AI systems that autonomously execute multi-step knowledge work across tools, queues and approval boundaries, owning outcomes end-to-end rather than assisting a human operator.","timeframes":{"optimistic":"by end of 2027","pragmatic":"2029-2031","sceptical":"mid-2030s or later"}},{"id":"scenario-robotics","definition":"General-purpose physical robots (humanoid or otherwise) that are commercially routine, not demo-quality, with fleet-deployable reliability and workable unit economics.","timeframes":{"optimistic":"by 2030","pragmatic":"2033-2036","sceptical":"not before the late 2030s"}},{"id":"scenario-software-automation","definition":"Coding agents reliably owning the loop from ticket to monitored production deploy across large codebases, leaving humans mostly specifying, reviewing and steering.","timeframes":{"optimistic":"by end of 2028","pragmatic":"2030-2032","sceptical":"not before the mid-2030s"}},{"id":"scenario-education-disruption","definition":"AI tutoring and assessment displacing institutional course and credential delivery as the primary structure through which people learn and signal mastery.","timeframes":{"optimistic":"by 2030","pragmatic":"2032-2035","sceptical":"not this generation"}}]},{"id":"review-2026-09-13","date":"2026-09-13","kind":"review","records":[{"id":"scenario-agi","definition":"Systems matching or exceeding human performance across most cognitive tasks, including ones outside their training distribution, without task-specific retraining.","timeframes":{"optimistic":"by end of 2028","pragmatic":"2032-2035","sceptical":"not before the 2040s"},"reason":"The threshold depends on the definition. We retain these windows as prompts for discussion. Neither a definition nor a benchmark result establishes a universal arrival year.","evidence":[{"title":"Google DeepMind: Levels of AGI","url":"https://deepmind.google/research/publications/66938/","date_label":"ICML 2024","finding":"This framework distinguishes breadth and depth of capability, with autonomy considered separately.","limit":"A framework for describing progress is not a certification that a current model has passed every threshold."},{"title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3","date_label":"2026 benchmark","finding":"Interactive environments test how agents explore, discover goals and adapt to unfamiliar tasks.","limit":"This measures a particular form of adaptation. A score alone cannot establish competence across all cognitive or economic work."}]},{"id":"scenario-agentic-work","definition":"AI systems that autonomously execute multi-step knowledge work across tools, queues and approval boundaries, owning outcomes end-to-end rather than assisting a human operator.","timeframes":{"optimistic":"by end of 2027","pragmatic":"2029-2031","sceptical":"mid-2030s or later"},"reason":"Task completion is measurable, but reliable ownership of a workplace process is a broader claim. The dates remain illustrative, with no new adoption probability assigned.","evidence":[{"title":"METR: task-completion time horizons","url":"https://metr.org/time-horizons/","date_label":"Page updated 8 May 2026","finding":"METR estimates task difficulty at specified success rates, using human completion time as the reference.","limit":"These are mainly well-specified software tasks. A time horizon is not continuous autonomous runtime or proof that a whole job can be automated."},{"title":"METR: metrics of agent ability","url":"https://metr.org/notes/2026-07-24-metrics-of-model-ability/","date_label":"24 July 2026","finding":"The research note compares performance measures that account for time, cost and the human baseline.","limit":"It analyses measurement choices. It does not establish widespread enterprise adoption."}]},{"id":"scenario-robotics","definition":"General-purpose physical robots (humanoid or otherwise) that are commercially routine, not demo-quality, with fleet-deployable reliability and workable unit economics.","timeframes":{"optimistic":"by 2030","pragmatic":"2033-2036","sceptical":"not before the late 2030s"},"reason":"The July model cards provide concrete evidence of robotics work and its limits. They do not establish fleet economics or commercial general-purpose reliability. The date windows remain illustrative.","evidence":[{"title":"Google DeepMind: Gemini Robotics ER 2","url":"https://deepmind.google/models/model-cards/gemini-robotics-er-2/","date_label":"30 July 2026","finding":"The model card describes spatial and temporal reasoning, tool orchestration and success detection for physical agents.","limit":"The publisher restricts safety-critical uses. Reasoning capability alone does not demonstrate reliable physical deployment."},{"title":"Google DeepMind: Robotics On-Device 2","url":"https://deepmind.google/models/model-cards/gemini-robotics-on-device-2/","date_label":"30 July 2026","finding":"The card describes local robot action generation and evaluations of manipulation across several robot platforms.","limit":"Access is limited to trusted testers. The card identifies limits on unfamiliar tasks and complex robot control."}]},{"id":"scenario-software-automation","definition":"Coding agents reliably owning the loop from ticket to monitored production deploy across large codebases, leaving humans mostly specifying, reviewing and steering.","timeframes":{"optimistic":"by end of 2028","pragmatic":"2030-2032","sceptical":"not before the mid-2030s"},"reason":"The evidence supports testing useful delegation, with careful measurement of the whole task. It does not establish an arrival year for autonomous production delivery. The original windows remain illustrative.","evidence":[{"title":"METR: revising its productivity experiment","url":"https://metr.org/blog/2026-02-24-uplift-update/","date_label":"24 February 2026","finding":"METR reports that selection effects make its later developer experiment an unreliable estimate of the current productivity effect.","limit":"This is a reason to improve measurement. It establishes neither universal slowdown nor a reliable universal speedup."},{"title":"METR: task-completion methodology","url":"https://metr.org/time-horizons/","date_label":"Page updated 8 May 2026","finding":"The published methodology makes the task set, human reference and success thresholds inspectable.","limit":"Successful benchmark tasks do not include every responsibility involved in maintaining a live software service."}]},{"id":"scenario-education-disruption","definition":"AI tutoring and assessment displacing institutional course and credential delivery as the primary structure through which people learn and signal mastery.","timeframes":{"optimistic":"by 2030","pragmatic":"2032-2035","sceptical":"not this generation"},"reason":"Better tutoring and replacement of institutional education are different claims. The pragmatic case describes improvement within institutions, so it is a partial transformation rather than fulfilment of the full replacement definition. The dates remain illustrative.","evidence":[{"title":"Kestin and colleagues: AI tutoring trial","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12179260/","date_label":"Scientific Reports, 2025","finding":"A randomised trial in undergraduate physics reported higher immediate learning gains with a specifically designed AI tutor than with the comparison lessons.","limit":"A study of selected lessons does not establish long-term learning across subjects, replacement of teachers or a change in credentials."}]}]}],"briefs":[{"id":"scenario-agi","question":"If intelligence became abundant, what would still make our business valuable?","measure":"Quality, total cost and human intervention for the same piece of work.","moves":{"optimistic":"Test one offering built around a delivered outcome. Keep human approval for consequential decisions and learn where judgement still matters.","pragmatic":"Choose one expert workflow. Measure AI assistance against the current process before redesigning the role around it.","sceptical":"Build an advantage from useful tools and better organisational knowledge. Choose work whose value does not depend on an AGI deadline."}},{"id":"scenario-agentic-work","question":"What would we redesign if the same team could handle ten times as many routine cases?","measure":"For each 100 cases, compare completed work, exception handling time, recovery effort and customer outcomes with the baseline.","moves":{"optimistic":"Trial one bounded queue with a named owner, clear approval limits and tested recovery. Measure completed outcomes across the whole process.","pragmatic":"Choose one repeatable queue. Establish a baseline, then test assistance and delegation in stages, including the awkward cases.","sceptical":"Make case owners faster with better context, drafting and triage. Keep responsibility clear and find which handovers cause the most rework."}},{"id":"scenario-robotics","question":"If routine physical work became easier to automate, where would we choose to operate?","measure":"Uptime, interventions, task variation, maintenance effort and safety performance.","moves":{"optimistic":"Map one repetitive physical task, its exceptions and its safety boundaries. Define a small supervised trial before committing to a wider deployment.","pragmatic":"Identify a structured task and the changes its surroundings would need. Compare the full operating effort with the current process.","sceptical":"Improve the process, sensing and existing automation first. Keep a clear test for when a general-purpose robot would add enough value."}},{"id":"scenario-software-automation","question":"If producing code were ten times easier, what would still stop us shipping useful products?","measure":"Accepted changes, review time, escaped defects and recovery effort.","moves":{"optimistic":"Test one path from a clear ticket to a reviewed release. Include monitoring and rollback so the experiment measures delivery, not just code generation.","pragmatic":"Improve acceptance tests, codebase context and review queues around one workflow. Track whether the entire delivery cycle gets better.","sceptical":"Use assistance where results are easy to inspect. Work on the requirements, coordination and maintenance problems that remain after code is written."}},{"id":"scenario-education-disruption","question":"If everyone had a patient personal tutor, how would we develop and recognise expertise?","measure":"Retained learning, transfer to unfamiliar problems and performance without assistance.","moves":{"optimistic":"Pilot personalised practice for one defined skill. Agree how independent mastery will be assessed before expanding the learning programme.","pragmatic":"Add tutoring and feedback to an existing learning programme. Compare retained learning and teacher or mentor effort with the current approach.","sceptical":"Help teachers and mentors target their time. Test whether better preparation, feedback and practice improve learning within the existing programme."}}],"revision":"793e5684d201cb0c225c293446e37715d7d3b6ee6279bd1f073cf148f9ee8cac"}