{"schema":1,"history":[{"id":"predictions-2026-09-15","date":"2026-09-15","reason":"We are publishing our own forecasts with explicit milestones and year-end deadlines. Earlier countdowns followed illustrative scenario windows. Those remain in the alternative-scenario record, not as earlier versions of these predictions. The dates below are our editorial judgements after reviewing these sources, not dates supplied by the sources.","records":[{"id":"scenario-agi","title":"General intelligence","milestone":"A general-purpose AI will independently demonstrate skilled-human performance across a broad set of unfamiliar cognitive tasks.","target_date":"2031-12-31","resolution":["Two evaluation teams independent of the model provider test the same frozen general-purpose system on independently created tasks. Each preregisters task-construction rules, scoring and human recruitment before testing, then publishes its protocol, human comparison, tools, time budgets and results.","Each evaluation covers language, mathematics, coding, factual synthesis, scientific reasoning, visual-spatial reasoning, planning, social reasoning, learning new rules and switching between tasks. Each domain contains at least 100 tasks created after the system was frozen. No task-specific retraining is allowed.","The system reaches or exceeds the median of relevant skilled adult participants in every domain under the same stated time allowance. This is our operational AGI threshold, not a claim that every definition of AGI has been satisfied."],"rationale":"We expect broader reasoning, memory and tool use to improve over several more development cycles. We put our date at the end of 2031 because unfamiliar tasks, consistency across domains and independent replication are a much higher bar than a strong benchmark score. This is our judgement about the remaining work, not a trend line fitted to the sources.","uncertainty":"Very uncertain timing. A generalisation breakthrough could bring this forward sharply, while persistent failures on unfamiliar tasks could put it beyond 2031. Our test is a declared working definition, not an industry standard.","earlier":"Independent teams reproduce skilled-human performance across newly created domains without task-specific tuning or undisclosed human help.","later":"Public benchmark gains fail to transfer to fresh tasks, or broad competence continues to depend on specialist retraining and extensive human correction.","evidence":[{"title":"Google DeepMind: Levels of AGI","url":"https://deepmind.google/research/publications/66938/","date_label":"21 July 2024","finding":"The framework separates the breadth of a system's abilities, its performance and its autonomy. This informs our use of a broad human comparison.","limit":"It defines a way to discuss capability. It does not establish that our threshold has been met or predict an arrival year."},{"title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3","date_label":"2026 benchmark; checked 15 September 2026","finding":"Interactive tasks examine learning unfamiliar rules, exploring environments and adapting through experience.","limit":"Adaptive reasoning is only part of general intelligence. This benchmark alone cannot resolve our broader forecast."}]},{"id":"scenario-agentic-work","title":"Agents at work","milestone":"AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.","target_date":"2028-12-31","resolution":["Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.","Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.","People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses."],"rationale":"Bounded workflow execution already exists. We expect the next two years to turn isolated deployments into repeatable operations across sectors, with permissions, recovery and ownership becoming part of the product. End-2028 is our forecast for that operational milestone.","uncertainty":"Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.","earlier":"Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.","later":"Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.","evidence":[{"title":"Salesforce and Air India: workflow expansion","url":"https://www.salesforce.com/in/news/press-releases/2026/09/15/air-india-accelerates-customer-service-transformation-with-agentforce/","date_label":"15 September 2026","finding":"The announcement describes multi-system refund handling, passenger-name changes and actions across several customer-service intents.","limit":"This is a vendor announcement. It does not publish the autonomous completion share or independently assessed failure rates required by our forecast."},{"title":"Anthropic: measuring agent autonomy","url":"https://www.anthropic.com/research/measuring-agent-autonomy","date_label":"18 February 2026","finding":"The provider's interaction analysis examines how tool-using agents operate alongside human involvement and safeguards.","limit":"Tool-call activity from one provider does not establish end-to-end workflow success or representative enterprise adoption."}]},{"id":"scenario-robotics","title":"Robotics","milestone":"Adaptable robots will perform several useful jobs routinely across factories and warehouses, with demonstrated operating economics.","target_date":"2031-12-31","resolution":["One commercially available robot platform using a shared general-purpose control model runs for six months at ten production sites belonging to at least three unrelated customers. It performs at least three materially different task families overall and at least two at each site. Customers demonstrate a new task learned from instructions or examples without writing task-specific control code.","Customer-confirmed reporting shows at least 95% of task cycles completed without human intervention. The reporting period, task mix, failures and supervised or teleoperated time are disclosed.","Customer-published cost per successful task matches or beats the incumbent process, including maintenance and supervision. This is a commercial-site milestone, not a prediction of routine household humanoids."],"rationale":"Useful paid robot work is already happening. Our end-2031 forecast allows time for broader task handling, several years of deployment and credible customer economics to follow today's more specialised systems. The remaining challenge is repeatability across sites and tasks.","uncertainty":"The date is uncertain because reliability, safety and manufacturing must improve together. Strong demonstrations or a product launch would not be enough to count this as achieved.","earlier":"Customers publish six-month multi-task fleet results, lower intervention rates and repeat purchases supported by full operating costs.","later":"Maintenance, low utilisation, teleoperation or poor transfer between tasks keeps the cost of successful work above the existing process.","evidence":[{"title":"BMW: Figure 03 production programme","url":"https://www.press.bmwgroup.com/global/article/detail/T0458778EN/bmw-group-advances-the-use-of-physical-ai-in-production-with-figure-03-project-in-spartanburg?language=en","date_label":"25 June 2026","finding":"The customer describes earlier humanoid work in production and a programme extending towards logistics sequencing.","limit":"A customer programme does not establish broad multi-site task coverage or the comparative costs in our resolution rules."},{"title":"Agility Robotics: Digit 5 announcement","url":"https://www.agilityrobotics.com/content/agility-unveils-digit-5-humanoid-robot-built-for-cooperatively-safe-work-at-scale","date_label":"15 September 2026","finding":"Agility reports substantial operational experience with Digit 4 and plans wider task capabilities for Digit 5, with availability expected by end-2027.","limit":"Operational figures are vendor-reported. Digit 5 capabilities and availability are forward-looking plans, not completed outcomes."}]},{"id":"scenario-software-automation","title":"Software creation","milestone":"Coding agents will take most routine software tickets from specification to monitored production in at least ten organisations, with human approval.","target_date":"2027-12-31","resolution":["Ten named organisations each publish customer-confirmed or independently evaluated results for at least 100 consecutive eligible tickets over at least 90 days. Eligibility is fixed before assignment: routine fixes or small features estimated at no more than one engineer-day.","More than half reach implementation, tests, review responses, approved deployment and seven days of monitoring without people writing or correcting the code. Humans still set requirements, review and authorise release.","Failed deployments, rollbacks and manually rescued tickets remain in the denominator. Merged pull requests alone do not meet the milestone. This does not predict autonomous ownership of all engineering work."],"rationale":"Substantial execution already works in well-prepared engineering environments. We expect reproducible routine delivery across ten organisations by end-2027, with deployment and monitoring included. Human approval makes this a nearer milestone than unsupervised engineering across arbitrary codebases.","uncertainty":"This is an assertive forecast. Repository preparation, review effort and access to credible post-deployment results could delay verification. The date is our judgement, not a provider's commitment.","earlier":"Organisations publish consecutive-ticket results with little manual rescue and reliable post-deployment checks.","later":"Review effort, rollback rates or legacy-system failures prevent agents reliably completing the whole delivery cycle.","evidence":[{"title":"OpenAI: harness engineering","url":"https://openai.com/index/harness-engineering/","date_label":"11 February 2026","finding":"An engineering team describes agents producing changes, verifying application behaviour, responding to feedback and progressing work to merge.","limit":"The results rely on that repository's structure and tooling. They do not prove general performance or autonomous production monitoring."},{"title":"Stripe: Minions coding agents","url":"https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents","date_label":"9 February 2026","finding":"Stripe reports a substantial weekly volume of agent-generated merged pull requests with human review.","limit":"Merged change volume does not reveal the eligible-ticket completion fraction, rescue rate or full production-delivery performance."},{"title":"METR: developer productivity experiment update","url":"https://metr.org/blog/2026-02-24-uplift-update/","date_label":"24 February 2026","finding":"The follow-up experiment describes selection effects that make the size of current productivity gains difficult to measure reliably.","limit":"A productivity experiment is not an end-to-end autonomy test. It cautions against generalising from selected demonstrations."}]},{"id":"scenario-education-disruption","title":"Education","milestone":"Generative AI tutoring will show repeatable learning gains that last after the AI is put away, across schools or colleges.","target_date":"2029-12-31","resolution":["Two independently evaluated, preregistered randomised trials of generative AI tutoring collectively cover at least 2,000 learners across at least 20 schools or colleges, with an intervention lasting at least one academic term. Each compares against a non-AI control with comparable teaching and learning time.","Both trials show statistically positive effects on unassisted assessments at least eight weeks after the intervention. Their combined standardised effect is at least 0.2, using an inverse-variance weighted mean of the reported primary outcomes.","Human teachers and safeguarding remain in place. The reports disclose attrition and assistance. This predicts demonstrated durable learning gains, not the replacement of teachers, schools or credentials."],"rationale":"Promising short trials and larger programmes are already under way, while retained learning remains less certain. We expect stronger designs, full-term studies and independent replication over the next three school years to meet this milestone by end-2029.","uncertainty":"Benefits may depend heavily on tutor design, subject and student group. Immediate performance with assistance can improve without producing durable independent learning.","earlier":"Multi-site trials show replicated gains on delayed, unassisted tests across learner groups.","later":"Gains disappear without AI, fail to transfer, or depend on levels of human support that do not scale.","evidence":[{"title":"Kestin and colleagues: AI tutoring trial","url":"https://www.nature.com/articles/s41598-025-97652-6","date_label":"3 June 2025","finding":"A randomised undergraduate physics study reports stronger immediate learning with a specifically designed AI tutor.","limit":"The short intervention at one institution does not demonstrate long-term retention or broad institutional scale."},{"title":"NBER: large-scale AI tutoring experiment","url":"https://www.nber.org/papers/w35621","date_label":"August 2026 working paper","finding":"A trial involving thousands of pupils reports improved recovery after mistakes but slower progress, with limited evidence of delayed learning gains.","limit":"The reported delayed assessment is much shorter than our eight-week threshold, and a working paper is not independent replication."}]}]},{"date":"2026-09-16","id":"predictions-2026-09-16","reason":"We reassessed all five predictions against their published tests. The reviewed deployments, evaluations and tutoring studies do not justify moving a target or changing a milestone. All five dates stay unchanged. This is an evidence review, not a claim that another day passing changes our forecast.","records":[{"earlier":"Independent teams reproduce skilled-human performance across newly created domains without task-specific tuning or undisclosed human help.","evidence":[{"date_label":"21 July 2024","finding":"The framework separates the breadth of a system's abilities, its performance and its autonomy. This informs our use of a broad human comparison.","limit":"It defines a way to discuss capability. It does not establish that our threshold has been met or predict an arrival year.","title":"Google DeepMind: Levels of AGI","url":"https://deepmind.google/research/publications/66938/"},{"date_label":"2026 benchmark, checked 16 September 2026","finding":"Interactive tasks examine learning unfamiliar rules, exploring environments and adapting through experience.","limit":"Adaptive reasoning is only part of general intelligence. This benchmark alone cannot resolve our broader forecast.","title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3"}],"id":"scenario-agi","later":"Public benchmark gains fail to transfer to fresh tasks, or broad competence continues to depend on specialist retraining and extensive human correction.","milestone":"A general-purpose AI will independently demonstrate skilled-human performance across a broad set of unfamiliar cognitive tasks.","rationale":"We retain end-2031. ARC-AGI-3 tests learning unfamiliar interactive tasks, while the Levels of AGI framework separates breadth, performance and autonomy. Neither establishes our two-team, ten-domain human comparison. These remain useful checks on generalisation, not evidence for an earlier arrival date. The timing is still our uncertain editorial judgement.","resolution":["Two evaluation teams independent of the model provider test the same frozen general-purpose system on independently created tasks. Each preregisters task-construction rules, scoring and human recruitment before testing, then publishes its protocol, human comparison, tools, time budgets and results.","Each evaluation covers language, mathematics, coding, factual synthesis, scientific reasoning, visual-spatial reasoning, planning, social reasoning, learning new rules and switching between tasks. Each domain contains at least 100 tasks created after the system was frozen. No task-specific retraining is allowed.","The system reaches or exceeds the median of relevant skilled adult participants in every domain under the same stated time allowance. This is our operational AGI threshold, not a claim that every definition of AGI has been satisfied."],"target_date":"2031-12-31","title":"General intelligence","uncertainty":"Very uncertain timing. A generalisation breakthrough could bring this forward sharply, while persistent failures on unfamiliar tasks could put it beyond 2031. Our test is a declared working definition, not an industry standard."},{"earlier":"Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.","evidence":[{"date_label":"15 September 2026","finding":"The announcement describes refund handling, passenger-name changes and multi-intent customer-service actions, with shorter reported turnaround times.","limit":"Its above-95% confidence threshold for email responses is not a measured workflow completion rate. The vendor report does not publish our consecutive-case denominator or autonomous success share.","title":"Salesforce and Air India: workflow expansion","url":"https://www.salesforce.com/in/news/press-releases/2026/09/15/air-india-accelerates-customer-service-transformation-with-agentforce/"},{"date_label":"18 February 2026","finding":"The provider observes longer high-end agent turns and studies how human oversight varies with experience and task risk.","limit":"Duration is not successful completion. The API analysis cannot reconstruct full sessions from isolated tool calls, and neither dataset establishes our ten-organisation workflow threshold.","title":"Anthropic: measuring agent autonomy","url":"https://www.anthropic.com/research/measuring-agent-autonomy"}],"id":"scenario-agentic-work","later":"Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.","milestone":"AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.","rationale":"We retain end-2028. Air India's expanded workflows strengthen the case that bounded agents can do useful operational work. The announcement still lacks consecutive-case completion and intervention rates across the organisations our test requires. Longer agent activity in Anthropic's study is not equivalent to completed business workflows. We have no sufficient verified basis for moving the date.","resolution":["Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.","Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.","People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses."],"target_date":"2028-12-31","title":"Agents at work","uncertainty":"Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it."},{"earlier":"Customers publish six-month multi-task fleet results, lower intervention rates and repeat purchases supported by full operating costs.","evidence":[{"date_label":"25 June 2026","finding":"The customer describes earlier humanoid work in production and a programme extending towards logistics sequencing.","limit":"A customer programme does not establish broad multi-site task coverage or the comparative costs in our resolution rules.","title":"BMW: Figure 03 production programme","url":"https://www.press.bmwgroup.com/global/article/detail/T0458778EN/bmw-group-advances-the-use-of-physical-ai-in-production-with-figure-03-project-in-spartanburg?language=en"},{"date_label":"15 September 2026","finding":"Agility reports substantial operational experience with Digit 4 and plans wider task capabilities for Digit 5, with availability expected by end-2027.","limit":"Operational figures are vendor-reported. Digit 5 capabilities and availability are forward-looking plans, not completed outcomes.","title":"Agility Robotics: Digit 5 announcement","url":"https://www.agilityrobotics.com/content/agility-unveils-digit-5-humanoid-robot-built-for-cooperatively-safe-work-at-scale"},{"title":"Agility Robotics: Digit 5 development status","url":"https://www.agilityrobotics.com/solutions/digit-5","date_label":"Product page, checked 16 September 2026","finding":"The product page explicitly describes Digit 5 as in development, with estimated specifications and some safety functions still being developed.","limit":"Product specifications and visualisations cannot establish sustained customer performance, comparative operating costs or delivery of the announced capabilities."}],"id":"scenario-robotics","later":"Maintenance, low utilisation, teleoperation or poor transfer between tasks keeps the cost of successful work above the existing process.","milestone":"Adaptable robots will perform several useful jobs routinely across factories and warehouses, with demonstrated operating economics.","rationale":"We retain end-2031. BMW reports production experience and a further logistics programme. Agility's 15 September announcement builds on Digit 4 experience, but Digit 5 remains a development programme with general availability expected by end-2027. These reports do not yet demonstrate our six-month, multi-task, ten-site economics test. A planned product is not a completed deployment.","resolution":["One commercially available robot platform using a shared general-purpose control model runs for six months at ten production sites belonging to at least three unrelated customers. It performs at least three materially different task families overall and at least two at each site. Customers demonstrate a new task learned from instructions or examples without writing task-specific control code.","Customer-confirmed reporting shows at least 95% of task cycles completed without human intervention. The reporting period, task mix, failures and supervised or teleoperated time are disclosed.","Customer-published cost per successful task matches or beats the incumbent process, including maintenance and supervision. This is a commercial-site milestone, not a prediction of routine household humanoids."],"target_date":"2031-12-31","title":"Robotics","uncertainty":"The date is uncertain because reliability, safety and manufacturing must improve together. Strong demonstrations or a product launch would not be enough to count this as achieved."},{"earlier":"Organisations publish consecutive-ticket results with little manual rescue and reliable post-deployment checks.","evidence":[{"date_label":"11 February 2026","finding":"An engineering team describes agents producing changes, verifying application behaviour, responding to feedback and progressing work to merge.","limit":"The results rely on that repository's structure and tooling. They do not prove general performance or autonomous production monitoring.","title":"OpenAI: harness engineering","url":"https://openai.com/index/harness-engineering/"},{"date_label":"9 February 2026","finding":"Stripe reports a substantial weekly volume of agent-generated merged pull requests with human review.","limit":"Merged change volume does not reveal the eligible-ticket completion fraction, rescue rate or full production-delivery performance.","title":"Stripe: Minions coding agents","url":"https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents"},{"date_label":"24 February 2026","finding":"METR says selection and time-measurement problems make its follow-up an unreliable estimate of current productivity gains. It is redesigning the experiment.","limit":"This is not evidence that agents cannot help, nor a measurement of autonomous delivery. It cannot resolve our production-ticket milestone.","title":"METR: developer productivity experiment update","url":"https://metr.org/blog/2026-02-24-uplift-update/"}],"id":"scenario-software-automation","later":"Review effort, rollback rates or legacy-system failures prevent agents reliably completing the whole delivery cycle.","milestone":"Coding agents will take most routine software tickets from specification to monitored production in at least ten organisations, with human approval.","rationale":"We retain end-2027, with high uncertainty. OpenAI's prepared engineering environment and Stripe's reported agent-generated changes show useful execution. They do not supply our full consecutive-ticket denominator through deployment and seven days of monitoring across ten organisations. METR's selection warnings also make simple productivity extrapolation unsafe. The reviewed evidence does not justify bringing the date forward.","resolution":["Ten named organisations each publish customer-confirmed or independently evaluated results for at least 100 consecutive eligible tickets over at least 90 days. Eligibility is fixed before assignment: routine fixes or small features estimated at no more than one engineer-day.","More than half reach implementation, tests, review responses, approved deployment and seven days of monitoring without people writing or correcting the code. Humans still set requirements, review and authorise release.","Failed deployments, rollbacks and manually rescued tickets remain in the denominator. Merged pull requests alone do not meet the milestone. This does not predict autonomous ownership of all engineering work."],"target_date":"2027-12-31","title":"Software creation","uncertainty":"This is an assertive forecast. Repository preparation, review effort and access to credible post-deployment results could delay verification. The date is our judgement, not a provider's commitment."},{"earlier":"Multi-site trials show replicated gains on delayed, unassisted tests across learner groups.","evidence":[{"date_label":"3 June 2025","finding":"A randomised undergraduate physics study reports stronger immediate learning with a specifically designed AI tutor.","limit":"The short intervention at one institution does not demonstrate long-term retention or broad institutional scale.","title":"Kestin and colleagues: AI tutoring trial","url":"https://www.nature.com/articles/s41598-025-97652-6"},{"title":"Oreopoulos and colleagues: mastery-based AI tutoring experiment","url":"https://edworkingpapers.com/ai26-1552","date_label":"August 2026 working paper","finding":"The authors report a randomised trial with over 6,000 pupils. AI helped recovery after mistakes, but delayed gains were limited and concentrated on practised material in the mastery workflow.","limit":"The delayed test was one week later, not eight weeks. The most encouraging gains were only marginally significant. This is the same study as NBER working paper 35621, not an independent replication."}],"id":"scenario-education-disruption","later":"Gains disappear without AI, fail to transfer, or depend on levels of human support that do not scale.","milestone":"Generative AI tutoring will show repeatable learning gains that last after the AI is put away, across schools or colleges.","rationale":"We retain end-2029. The short physics trial supports a carefully designed tutor, while the larger August maths experiment finds its most encouraging delayed results in a structured mastery workflow. Its one-week assessment remains far short of our eight-week retention test. We still need full-term studies and independent replication, not just improved answers during practice.","resolution":["Two independently evaluated, preregistered randomised trials of generative AI tutoring collectively cover at least 2,000 learners across at least 20 schools or colleges, with an intervention lasting at least one academic term. Each compares against a non-AI control with comparable teaching and learning time.","Both trials show statistically positive effects on unassisted assessments at least eight weeks after the intervention. Their combined standardised effect is at least 0.2, using an inverse-variance weighted mean of the reported primary outcomes.","Human teachers and safeguarding remain in place. The reports disclose attrition and assistance. This predicts demonstrated durable learning gains, not the replacement of teachers, schools or credentials."],"target_date":"2029-12-31","title":"Education","uncertainty":"Benefits may depend heavily on tutor design, subject and student group. Immediate performance with assistance can improve without producing durable independent learning."}]},{"date":"2026-09-18","id":"predictions-2026-09-18","reason":"We reassessed all five milestones against primary material on 18 September. We retain every target and resolution rule. The evidence supports useful bounded capabilities, but does not supply the breadth, consecutive-case denominators, operating economics or delayed learning results our tests require. This is a limited source review, not a claim that all relevant research was surveyed. No date moves because time has passed.","records":[{"earlier":"Independent teams reproduce skilled-human performance across newly created domains without task-specific tuning or undisclosed human help.","evidence":[{"date_label":"21 July 2024","finding":"The framework separates the breadth of a system's abilities, its performance and its autonomy. This informs our use of a broad human comparison.","limit":"It defines a way to discuss capability. It does not establish that our threshold has been met or predict an arrival year.","title":"Google DeepMind: Levels of AGI","url":"https://deepmind.google/research/publications/66938/"},{"date_label":"Benchmark definition, checked 18 September 2026","finding":"ARC-AGI-3 asks agents to discover goals and adapt in unfamiliar interactive environments. Its score compares game-solving efficiency with human performance.","limit":"The benchmark definition is not a claim that a model has met our broad threshold. We did not verify a new leaderboard result in this review.","title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3"}],"id":"scenario-agi","later":"Public benchmark gains fail to transfer to fresh tasks, or broad competence continues to depend on specialist retraining and extensive human correction.","milestone":"A general-purpose AI will independently demonstrate skilled-human performance across a broad set of unfamiliar cognitive tasks.","rationale":"We retain end-2031 with very high uncertainty. ARC-AGI-3 tests adaptation to unfamiliar interactive environments and compares efficiency with humans. That is relevant progress measurement, but it is not our preregistered, two-team comparison across ten cognitive domains. The Levels of AGI paper also separates breadth, performance and autonomy. These sources do not provide a new timing basis for changing our judgement.","resolution":["Two evaluation teams independent of the model provider test the same frozen general-purpose system on independently created tasks. Each preregisters task-construction rules, scoring and human recruitment before testing, then publishes its protocol, human comparison, tools, time budgets and results.","Each evaluation covers language, mathematics, coding, factual synthesis, scientific reasoning, visual-spatial reasoning, planning, social reasoning, learning new rules and switching between tasks. Each domain contains at least 100 tasks created after the system was frozen. No task-specific retraining is allowed.","The system reaches or exceeds the median of relevant skilled adult participants in every domain under the same stated time allowance. This is our operational AGI threshold, not a claim that every definition of AGI has been satisfied."],"target_date":"2031-12-31","title":"General intelligence","uncertainty":"Very uncertain timing. A generalisation breakthrough could bring this forward sharply, while persistent failures on unfamiliar tasks could put it beyond 2031. Our test is a declared working definition, not an industry standard."},{"earlier":"Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.","evidence":[{"date_label":"15 September 2026","finding":"The announcement describes refund handling, passenger-name changes and multi-intent customer-service actions, with shorter reported turnaround times.","limit":"Its above-95% confidence threshold for email responses is not a measured workflow completion rate. The vendor report does not publish our consecutive-case denominator or autonomous success share.","title":"Salesforce and Air India: workflow expansion","url":"https://www.salesforce.com/in/news/press-releases/2026/09/15/air-india-accelerates-customer-service-transformation-with-agentforce/"},{"date_label":"18 February 2026","finding":"The provider observes longer high-end agent turns and studies how human oversight varies with experience and task risk.","limit":"Duration is not successful completion. The API analysis cannot reconstruct full sessions from isolated tool calls, and neither dataset establishes our ten-organisation workflow threshold.","title":"Anthropic: measuring agent autonomy","url":"https://www.anthropic.com/research/measuring-agent-autonomy"}],"id":"scenario-agentic-work","later":"Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.","milestone":"AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.","rationale":"We retain end-2028. Air India's customer-endorsed vendor announcement describes useful multi-system action, but its response-confidence threshold is not an observed autonomous completion rate. Anthropic's tool-call analysis cannot reconstruct complete API sessions. Neither supplies ten organisations' consecutive-case results over 90 days. These gaps do not disprove progress, but they leave insufficient evidence to revise our date.","resolution":["Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.","Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.","People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses."],"target_date":"2028-12-31","title":"Agents at work","uncertainty":"Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it."},{"earlier":"Customers publish six-month multi-task fleet results, lower intervention rates and repeat purchases supported by full operating costs.","evidence":[{"date_label":"15 September 2026","finding":"Agility reports more than 65,000 Digit 4 operating hours and approximately 98% accuracy while on-task at GXO's Flowery Branch site. Digit 5 general availability is expected by end-2027.","limit":"These are vendor-reported figures. On-task accuracy is not an all-cycle intervention-free success rate, and the announcement does not supply our ten-site task mix or comparative cost per successful task.","title":"Agility Robotics: Digit 5 announcement","url":"https://www.agilityrobotics.com/content/agility-unveils-digit-5-humanoid-robot-built-for-cooperatively-safe-work-at-scale"},{"title":"Agility Robotics: Digit 5 development status","url":"https://www.agilityrobotics.com/solutions/digit-5","date_label":"Product page, checked 18 September 2026","finding":"The product page explicitly describes Digit 5 as in development, with estimated specifications and some safety functions still being developed.","limit":"Product specifications and visualisations cannot establish sustained customer performance, comparative operating costs or delivery of the announced capabilities."}],"id":"scenario-robotics","later":"Maintenance, low utilisation, teleoperation or poor transfer between tasks keeps the cost of successful work above the existing process.","milestone":"Adaptable robots will perform several useful jobs routinely across factories and warehouses, with demonstrated operating economics.","rationale":"We retain end-2031. Agility's reported Digit 4 operating experience supports the case for useful paid work. Its approximately 98% accuracy while on-task is not our intervention-free success measure across ten sites, and its announcement lacks comparable cost per successful task. Digit 5 remains in development with general availability expected by end-2027. We cannot infer six-month, multi-task economics from those figures.","resolution":["One commercially available robot platform using a shared general-purpose control model runs for six months at ten production sites belonging to at least three unrelated customers. It performs at least three materially different task families overall and at least two at each site. Customers demonstrate a new task learned from instructions or examples without writing task-specific control code.","Customer-confirmed reporting shows at least 95% of task cycles completed without human intervention. The reporting period, task mix, failures and supervised or teleoperated time are disclosed.","Customer-published cost per successful task matches or beats the incumbent process, including maintenance and supervision. This is a commercial-site milestone, not a prediction of routine household humanoids."],"target_date":"2031-12-31","title":"Robotics","uncertainty":"The date is uncertain because reliability, safety and manufacturing must improve together. Strong demonstrations or a product launch would not be enough to count this as achieved."},{"earlier":"Organisations publish consecutive-ticket results with little manual rescue and reliable post-deployment checks.","evidence":[{"date_label":"11 February 2026","finding":"An engineering team describes agents producing changes, verifying application behaviour, responding to feedback and progressing work to merge.","limit":"The results rely on that repository's structure and tooling. They do not prove general performance or autonomous production monitoring.","title":"OpenAI: harness engineering","url":"https://openai.com/index/harness-engineering/"},{"date_label":"24 February 2026","finding":"METR says selection and time-measurement problems make its follow-up an unreliable estimate of current productivity gains. It is redesigning the experiment.","limit":"This is not evidence that agents cannot help, nor a measurement of autonomous delivery. It cannot resolve our production-ticket milestone.","title":"METR: developer productivity experiment update","url":"https://metr.org/blog/2026-02-24-uplift-update/"}],"id":"scenario-software-automation","later":"Review effort, rollback rates or legacy-system failures prevent agents reliably completing the whole delivery cycle.","milestone":"Coding agents will take most routine software tickets from specification to monitored production in at least ten organisations, with human approval.","rationale":"We retain end-2027 as an assertive forecast. OpenAI reports agent-written delivery in a deliberately prepared repository, including observability and review tooling. METR's follow-up cautions that selection and time-measurement problems obstruct reliable productivity estimates. Useful execution is credible, but neither report measures our consecutive eligible tickets through seven days of production monitoring across ten organisations. We do not use this site's single Feral cycle as population evidence.","resolution":["Ten named organisations each publish customer-confirmed or independently evaluated results for at least 100 consecutive eligible tickets over at least 90 days. Eligibility is fixed before assignment: routine fixes or small features estimated at no more than one engineer-day.","More than half reach implementation, tests, review responses, approved deployment and seven days of monitoring without people writing or correcting the code. Humans still set requirements, review and authorise release.","Failed deployments, rollbacks and manually rescued tickets remain in the denominator. Merged pull requests alone do not meet the milestone. This does not predict autonomous ownership of all engineering work."],"target_date":"2027-12-31","title":"Software creation","uncertainty":"This is an assertive forecast. Repository preparation, review effort and access to credible post-deployment results could delay verification. The date is our judgement, not a provider's commitment."},{"earlier":"Multi-site trials show replicated gains on delayed, unassisted tests across learner groups.","evidence":[{"title":"Oreopoulos and colleagues: mastery-based AI tutoring experiment","url":"https://edworkingpapers.com/ai26-1552","date_label":"August 2026 working paper","finding":"The authors report a randomised trial with over 6,000 pupils. AI helped recovery after mistakes, but delayed gains were limited and concentrated on practised material in the mastery workflow.","limit":"The delayed test was one week later, not eight weeks. The most encouraging gains were only marginally significant. This is the same study as NBER working paper 35621, not an independent replication."}],"id":"scenario-education-disruption","later":"Gains disappear without AI, fail to transfer, or depend on levels of human support that do not scale.","milestone":"Generative AI tutoring will show repeatable learning gains that last after the AI is put away, across schools or colleges.","rationale":"We retain end-2029 with substantial uncertainty. In the August maths experiment, structured AI support improved recovery after mistakes, but the encouraging delayed gains were marginal and concentrated on practised material. The delayed test was one week later. That falls short of our eight-week unassisted assessment and independent replication requirements. This limited review adds no verified result that warrants accelerating or delaying the forecast.","resolution":["Two independently evaluated, preregistered randomised trials of generative AI tutoring collectively cover at least 2,000 learners across at least 20 schools or colleges, with an intervention lasting at least one academic term. Each compares against a non-AI control with comparable teaching and learning time.","Both trials show statistically positive effects on unassisted assessments at least eight weeks after the intervention. Their combined standardised effect is at least 0.2, using an inverse-variance weighted mean of the reported primary outcomes.","Human teachers and safeguarding remain in place. The reports disclose attrition and assistance. This predicts demonstrated durable learning gains, not the replacement of teachers, schools or credentials."],"target_date":"2029-12-31","title":"Education","uncertainty":"Benefits may depend heavily on tutor design, subject and student group. Immediate performance with assistance can improve without producing durable independent learning."}]},{"date":"2026-09-21","id":"predictions-2026-09-21","reason":"We reassessed all five milestones against primary material retrieved on 21 September. Long-horizon formal mathematics, high-volume customer workflows and cross-environment robot generalisation strengthen the evidence that bounded systems are improving. They still do not provide the independent breadth, consecutive eligible-case results, production-ticket outcomes, commercial robot economics or delayed learning replication in our tests. We therefore retain every target and resolution rule, with zero days added or removed.","records":[{"earlier":"Independent teams reproduce skilled-human performance across newly created domains without task-specific tuning or undisclosed human help.","evidence":[{"title":"Anthropic: formalising Fermat's Last Theorem","url":"https://www.anthropic.com/research/formalizing-fermats-last-theorem","date_label":"4 September 2026","finding":"Anthropic reports that Claude worked largely autonomously for 11 days to produce a computer-checked formalisation spanning 13 million lines of Lean and 29,500 intermediate theorems used in the final proof.","limit":"This formalises a known proof in one specialist domain with occasional high-level human direction. It is provider-reported evidence of long-horizon work, not our independent two-team comparison across ten unfamiliar cognitive domains."},{"title":"ARC Prize: ARC-AGI-3","url":"https://arcprize.org/arc-agi/3","date_label":"Benchmark definition, checked 21 September 2026","finding":"ARC-AGI-3 tests experience-driven adaptation, long-horizon planning and skill acquisition in novel interactive environments.","limit":"The published benchmark definition does not establish that one frozen system has met our broad human threshold. We did not verify a qualifying independent result."}],"id":"scenario-agi","later":"Public benchmark gains fail to transfer to fresh tasks, or broad competence continues to depend on specialist retraining and extensive human correction.","milestone":"A general-purpose AI will independently demonstrate skilled-human performance across a broad set of unfamiliar cognitive tasks.","rationale":"We retain end-2031 with very high uncertainty. The computer-checked Fermat formalisation is striking evidence of sustained execution in a specialist domain, while ARC-AGI-3 tests adaptation in unfamiliar interactive environments. Neither supplies two independent, preregistered evaluations of the same frozen system across our ten domains. The new evidence raises our view of bounded long-horizon capability, but does not give a defensible basis for moving the broad forecast.","resolution":["Two evaluation teams independent of the model provider test the same frozen general-purpose system on independently created tasks. Each preregisters task-construction rules, scoring and human recruitment before testing, then publishes its protocol, human comparison, tools, time budgets and results.","Each evaluation covers language, mathematics, coding, factual synthesis, scientific reasoning, visual-spatial reasoning, planning, social reasoning, learning new rules and switching between tasks. Each domain contains at least 100 tasks created after the system was frozen. No task-specific retraining is allowed.","The system reaches or exceeds the median of relevant skilled adult participants in every domain under the same stated time allowance. This is our operational AGI threshold, not a claim that every definition of AGI has been satisfied."],"target_date":"2031-12-31","title":"General intelligence","uncertainty":"Very uncertain timing. A generalisation breakthrough could bring this forward sharply, while persistent failures on unfamiliar tasks could put it beyond 2031. Our test is a declared working definition, not an industry standard."},{"earlier":"Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.","evidence":[{"title":"OpenAI and Wayfair: supplier workflow automation","url":"https://openai.com/index/wayfair/","date_label":"11 March 2026; checked 21 September 2026","finding":"The customer story reports 41,000 supplier support tickets automated per month and model use in core supplier and catalogue workflows.","limit":"It does not publish a predeclared consecutive eligible-case denominator, autonomous completion share, manual-rescue count or 90-day acceptance results. It is a provider-hosted customer story for one organisation."},{"title":"OpenAI and Circles: CareX customer support","url":"https://openai.com/index/circles/","date_label":"3 August 2026; checked 21 September 2026","finding":"The customer story reports a 65% autonomous resolution rate across supported workflows, with specialist agents and human escalation.","limit":"The rate is below our 80% threshold and the story does not disclose 1,000 consecutive eligible cases, a 90-day period or results for ten named organisations. Supported-workflow selection may narrow the denominator."}],"id":"scenario-agentic-work","later":"Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.","milestone":"AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.","rationale":"We retain end-2028. Wayfair's reported monthly volume and Circles' reported autonomous resolution rate are stronger operational signals than activity or response confidence alone. They still lack our predeclared consecutive-case denominator, rescue accounting, 90-day window and ten-organisation coverage. The evidence moves the judgement from isolated examples towards scaled bounded workflows, but not enough to change the deadline.","resolution":["Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.","Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.","People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses."],"target_date":"2028-12-31","title":"Agents at work","uncertainty":"Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it."},{"earlier":"Customers publish six-month multi-task fleet results, lower intervention rates and repeat purchases supported by full operating costs.","evidence":[{"title":"Figure: Helix 2.5 zero-shot generalisation","url":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization","date_label":"17 September 2026","finding":"Figure reports one pretrained base model adapted to three whole-body behaviours and evaluated across 30 unseen homes. Its blind evaluation reports 56% zero-shot task success, versus 9% without Index pretraining.","limit":"The vendor-run home evaluation is not a six-month commercial deployment. Its 56% result is well below our 95% intervention-free threshold and it supplies no multi-customer operating cost per successful task."},{"title":"Figure: Figure 03 at BMW","url":"https://www.figure.ai/news/f-03-at-bmw","date_label":"30 June 2026; checked 21 September 2026","finding":"Figure describes a new sequencing demonstration at BMW after its earlier robot contributed to production of 30,000 vehicles.","limit":"A demonstration and earlier single-task deployment do not establish three task families, ten production sites, customer-confirmed intervention rates or comparative economics."}],"id":"scenario-robotics","later":"Maintenance, low utilisation, teleoperation or poor transfer between tasks keeps the cost of successful work above the existing process.","milestone":"Adaptable robots will perform several useful jobs routinely across factories and warehouses, with demonstrated operating economics.","rationale":"We retain end-2031. Helix 2.5 materially improves the evidence for transfer across unfamiliar environments: one base model supported three behaviours in 30 unseen homes. The reported 56% end-to-end success also makes the remaining reliability gap visible. These were vendor-run home trials, not six-month paid operations across customers, and no comparative cost per successful task was published. This is meaningful capability evidence without the commercial proof needed to move our date.","resolution":["One commercially available robot platform using a shared general-purpose control model runs for six months at ten production sites belonging to at least three unrelated customers. It performs at least three materially different task families overall and at least two at each site. Customers demonstrate a new task learned from instructions or examples without writing task-specific control code.","Customer-confirmed reporting shows at least 95% of task cycles completed without human intervention. The reporting period, task mix, failures and supervised or teleoperated time are disclosed.","Customer-published cost per successful task matches or beats the incumbent process, including maintenance and supervision. This is a commercial-site milestone, not a prediction of routine household humanoids."],"target_date":"2031-12-31","title":"Robotics","uncertainty":"The date is uncertain because reliability, safety and manufacturing must improve together. Strong demonstrations or a product launch would not be enough to count this as achieved."},{"earlier":"Organisations publish consecutive-ticket results with little manual rescue and reliable post-deployment checks.","evidence":[{"title":"OpenAI: harness engineering","url":"https://openai.com/index/harness-engineering/","date_label":"11 February 2026; reassessed 21 September 2026","finding":"An engineering team describes agents producing changes, checking application behaviour, responding to feedback and progressing work towards merge in a deliberately prepared repository.","limit":"The account does not publish 100 consecutive eligible tickets across ten organisations or follow them through approved deployment and seven days of monitoring."},{"title":"OpenAI and Circles: Codex development efficiency","url":"https://openai.com/index/circles/","date_label":"3 August 2026; checked 21 September 2026","finding":"Circles reports a 29% increase in development efficiency across design, coding assistance and unit testing, while specialists retain validation and final review.","limit":"A productivity figure is not the share of eligible tickets completed autonomously. The story does not report deployment, rollback, manual correction or seven-day monitoring outcomes."}],"id":"scenario-software-automation","later":"Review effort, rollback rates or legacy-system failures prevent agents reliably completing the whole delivery cycle.","milestone":"Coding agents will take most routine software tickets from specification to monitored production in at least ten organisations, with human approval.","rationale":"We retain end-2027 as an assertive forecast. Prepared agent-first repositories and Circles' reported development-efficiency gain support useful execution across design, code and tests. Neither source measures the proportion of predeclared routine tickets completed without human code correction through deployment and seven days of monitoring. We still lack the ten-organisation consecutive-ticket evidence required to move the forecast.","resolution":["Ten named organisations each publish customer-confirmed or independently evaluated results for at least 100 consecutive eligible tickets over at least 90 days. Eligibility is fixed before assignment: routine fixes or small features estimated at no more than one engineer-day.","More than half reach implementation, tests, review responses, approved deployment and seven days of monitoring without people writing or correcting the code. Humans still set requirements, review and authorise release.","Failed deployments, rollbacks and manually rescued tickets remain in the denominator. Merged pull requests alone do not meet the milestone. This does not predict autonomous ownership of all engineering work."],"target_date":"2027-12-31","title":"Software creation","uncertainty":"This is an assertive forecast. Repository preparation, review effort and access to credible post-deployment results could delay verification. The date is our judgement, not a provider's commitment."},{"earlier":"Multi-site trials show replicated gains on delayed, unassisted tests across learner groups.","evidence":[{"title":"Oreopoulos and colleagues: mastery-based AI tutoring experiment","url":"https://edworkingpapers.com/ai26-1552","date_label":"August 2026 working paper; reassessed 21 September 2026","finding":"The randomised field experiment covers more than 6,000 middle-school pupils. Structured AI support improved recovery after mistakes, while the clearest delayed gains were marginal and concentrated on practised material.","limit":"The delayed assessment was one week later, not eight weeks. The intervention was not a full academic term and this remains one study rather than two independent replications."}],"id":"scenario-education-disruption","later":"Gains disappear without AI, fail to transfer, or depend on levels of human support that do not scale.","milestone":"Generative AI tutoring will show repeatable learning gains that last after the AI is put away, across schools or colleges.","rationale":"We retain end-2029 with substantial uncertainty. Rechecking the large maths experiment confirms that it is relevant scale evidence, but not evidence of our required durability. Its delayed assessment was one week later, the encouraging gains were marginal and concentrated on practised material, and there is no independent term-long replication. We found no retrieved result that warrants moving the forecast.","resolution":["Two independently evaluated, preregistered randomised trials of generative AI tutoring collectively cover at least 2,000 learners across at least 20 schools or colleges, with an intervention lasting at least one academic term. Each compares against a non-AI control with comparable teaching and learning time.","Both trials show statistically positive effects on unassisted assessments at least eight weeks after the intervention. Their combined standardised effect is at least 0.2, using an inverse-variance weighted mean of the reported primary outcomes.","Human teachers and safeguarding remain in place. The reports disclose attrition and assistance. This predicts demonstrated durable learning gains, not the replacement of teachers, schools or credentials."],"target_date":"2029-12-31","title":"Education","uncertainty":"Benefits may depend heavily on tutor design, subject and student group. Immediate performance with assistance can improve without producing durable independent learning."}]}],"revision":"8bc84471b49d174aafef24dfae180007c565dfea6d59666537641d554ff78f9b"}