SOFT CAT.ai
FIND SOMETHING USEFUL

Horizon / The SOFT CAT .ai forecast

Our predictions.
Five countdowns.

What we expect. When we expect it.
Five calls on AI’s future, with the evidence behind them and a clear test of whether we were right.

Five predictions. Five day counts.

Days remaining to each milestone, recalculated when our prediction changes.

Compare alternative scenarios ↗Read selected prediction: Agents at work ↓

Snapshot 22 Sept 2026 · UTC

What changed the estimate? ↓

Time removes days. Evidence can change the estimate. Each count follows its own prediction. A review can add or remove days for that future, with the reason recorded below. Forecasts with the same target date share a day count. How we calculate the days ↓

Read our agents at work prediction ↓

Published predictions refresh while this page is open.

Our prediction / Agents at work

By end of 2028,
we expect this.

AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.

Deadline: · Reviewed

What sets this day count?

831 days until our prediction deadline. We subtract the current time from this prediction's target date.

No days added or removed. This is the effect of the latest evidence review, separate from days passing.

Why this estimate?

We retain end-2028. Wayfair's reported monthly volume and Circles' reported autonomous resolution rate are stronger operational signals than activity or response confidence alone. They still lack our predeclared consecutive-case denominator, rescue accounting, 90-day window and ten-organisation coverage. The evidence moves the judgement from isolated examples towards scaled bounded workflows, but not enough to change the deadline.

Where we could be wrong

Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.

What would count as this prediction coming true?

These are our published resolution criteria. We will assess the evidence against them and record the outcome.

  1. Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.
  2. Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.
  3. People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses.

What would move it earlier

Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.

What would move it later

Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.

The evidence behind our call

Agents at work: what we know.

Reviewed 21 Sept 2026

These sources inform our judgement. They do not determine the target date. We explain what each supports and what it leaves unanswered.

11 March 2026; checked 21 September 2026

OpenAI and Wayfair: supplier workflow automation ↗

The customer story reports 41,000 supplier support tickets automated per month and model use in core supplier and catalogue workflows.

The limitIt does not publish a predeclared consecutive eligible-case denominator, autonomous completion share, manual-rescue count or 90-day acceptance results. It is a provider-hosted customer story for one organisation.

3 August 2026; checked 21 September 2026

OpenAI and Circles: CareX customer support ↗

The customer story reports a 65% autonomous resolution rate across supported workflows, with specialist agents and human escalation.

The limitThe rate is below our 80% threshold and the story does not disclose 1,000 consecutive eligible cases, a 90-day period or results for ten named organisations. Supported-workflow selection may narrow the denominator.

Accountable to our earlier calls

Prediction history.

Agents at work

Every published prediction stays in the record, including its original date, milestone and evidence. A revision adds a new entry.

No days added or removedBy end of 2028Read ↓Close ↑

We reassessed all five milestones against primary material retrieved on 21 September. Long-horizon formal mathematics, high-volume customer workflows and cross-environment robot generalisation strengthen the evidence that bounded systems are improving. They still do not provide the independent breadth, consecutive eligible-case results, production-ticket outcomes, commercial robot economics or delayed learning replication in our tests. We therefore retain every target and resolution rule, with zero days added or removed.

Change at this review: No days added or removed. Passing days are not included in this change.

Our prediction: AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.

Deadline: 31 Dec 2028

What would count

  1. Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.
  2. Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.
  3. People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses.

Why this estimate: We retain end-2028. Wayfair's reported monthly volume and Circles' reported autonomous resolution rate are stronger operational signals than activity or response confidence alone. They still lack our predeclared consecutive-case denominator, rescue accounting, 90-day window and ten-organisation coverage. The evidence moves the judgement from isolated examples towards scaled bounded workflows, but not enough to change the deadline.

Uncertainty: Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.

Earlier: Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.

Later: Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.

11 March 2026; checked 21 September 2026

OpenAI and Wayfair: supplier workflow automation ↗

The customer story reports 41,000 supplier support tickets automated per month and model use in core supplier and catalogue workflows.

The limitIt does not publish a predeclared consecutive eligible-case denominator, autonomous completion share, manual-rescue count or 90-day acceptance results. It is a provider-hosted customer story for one organisation.

3 August 2026; checked 21 September 2026

OpenAI and Circles: CareX customer support ↗

The customer story reports a 65% autonomous resolution rate across supported workflows, with specialist agents and human escalation.

The limitThe rate is below our 80% threshold and the story does not disclose 1,000 consecutive eligible cases, a 90-day period or results for ten named organisations. Supported-workflow selection may narrow the denominator.

No days added or removedBy end of 2028Read ↓Close ↑

We reassessed all five milestones against primary material on 18 September. We retain every target and resolution rule. The evidence supports useful bounded capabilities, but does not supply the breadth, consecutive-case denominators, operating economics or delayed learning results our tests require. This is a limited source review, not a claim that all relevant research was surveyed. No date moves because time has passed.

Change at this review: No days added or removed. Passing days are not included in this change.

Our prediction: AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.

Deadline: 31 Dec 2028

What would count

  1. Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.
  2. Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.
  3. People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses.

Why this estimate: We retain end-2028. Air India's customer-endorsed vendor announcement describes useful multi-system action, but its response-confidence threshold is not an observed autonomous completion rate. Anthropic's tool-call analysis cannot reconstruct complete API sessions. Neither supplies ten organisations' consecutive-case results over 90 days. These gaps do not disprove progress, but they leave insufficient evidence to revise our date.

Uncertainty: Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.

Earlier: Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.

Later: Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.

15 September 2026

Salesforce and Air India: workflow expansion ↗

The announcement describes refund handling, passenger-name changes and multi-intent customer-service actions, with shorter reported turnaround times.

The limitIts above-95% confidence threshold for email responses is not a measured workflow completion rate. The vendor report does not publish our consecutive-case denominator or autonomous success share.

18 February 2026

Anthropic: measuring agent autonomy ↗

The provider observes longer high-end agent turns and studies how human oversight varies with experience and task risk.

The limitDuration is not successful completion. The API analysis cannot reconstruct full sessions from isolated tool calls, and neither dataset establishes our ten-organisation workflow threshold.

No days added or removedBy end of 2028Read ↓Close ↑

We reassessed all five predictions against their published tests. The reviewed deployments, evaluations and tutoring studies do not justify moving a target or changing a milestone. All five dates stay unchanged. This is an evidence review, not a claim that another day passing changes our forecast.

Change at this review: No days added or removed. Passing days are not included in this change.

Our prediction: AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.

Deadline: 31 Dec 2028

What would count

  1. Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.
  2. Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.
  3. People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses.

Why this estimate: We retain end-2028. Air India's expanded workflows strengthen the case that bounded agents can do useful operational work. The announcement still lacks consecutive-case completion and intervention rates across the organisations our test requires. Longer agent activity in Anthropic's study is not equivalent to completed business workflows. We have no sufficient verified basis for moving the date.

Uncertainty: Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.

Earlier: Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.

Later: Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.

15 September 2026

Salesforce and Air India: workflow expansion ↗

The announcement describes refund handling, passenger-name changes and multi-intent customer-service actions, with shorter reported turnaround times.

The limitIts above-95% confidence threshold for email responses is not a measured workflow completion rate. The vendor report does not publish our consecutive-case denominator or autonomous success share.

18 February 2026

Anthropic: measuring agent autonomy ↗

The provider observes longer high-end agent turns and studies how human oversight varies with experience and task risk.

The limitDuration is not successful completion. The API analysis cannot reconstruct full sessions from isolated tool calls, and neither dataset establishes our ten-organisation workflow threshold.

Initial estimateBy end of 2028Read ↓Close ↑

We are publishing our own forecasts with explicit milestones and year-end deadlines. Earlier countdowns followed illustrative scenario windows. Those remain in the alternative-scenario record, not as earlier versions of these predictions. The dates below are our editorial judgements after reviewing these sources, not dates supplied by the sources.

Our prediction: AI agents will run repeatable business workflows across at least ten organisations, with people managing exceptions.

Deadline: 31 Dec 2028

What would count

  1. Customer-confirmed or independently evaluated public reports name ten organisations across at least three sectors. Each runs two distinct production workflows spanning at least two business systems.
  2. Each workflow reports at least 1,000 consecutive eligible cases over at least 90 days. At least 80% finish without a person executing workflow steps. Eligibility, acceptance tests and escalation rules are fixed before the reporting period.
  3. People may authorise consequential decisions and handle exceptions. Reports include failures and escalations in the denominator. This predicts repeatable operation at ten organisations, not majority adoption or unattended businesses.

Why this estimate: Bounded workflow execution already exists. We expect the next two years to turn isolated deployments into repeatable operations across sectors, with permissions, recovery and ownership becoming part of the product. End-2028 is our forecast for that operational milestone.

Uncertainty: Timing remains uncertain. Vendor case studies rarely publish consecutive-case failure and intervention data, so the milestone could occur before enough public evidence exists to verify it.

Earlier: Customer-confirmed 90-day results show high autonomous completion across multiple systems and sectors, including failures and recovery costs.

Later: Manual rescue, fragile integrations or operating costs prevent teams expanding beyond isolated queues.

15 September 2026

Salesforce and Air India: workflow expansion ↗

The announcement describes multi-system refund handling, passenger-name changes and actions across several customer-service intents.

The limitThis is a vendor announcement. It does not publish the autonomous completion share or independently assessed failure rates required by our forecast.

18 February 2026

Anthropic: measuring agent autonomy ↗

The provider's interaction analysis examines how tool-using agents operate alongside human involvement and safeguards.

The limitTool-call activity from one provider does not establish end-to-end workflow success or representative enterprise adoption.

The question behind the timeline

Is AGI
already here?

Start with what you mean by general intelligence. The definition changes the evidence you need.

Our reading of the sources, checked . Explore the criteria before drawing a conclusion.

How general, and how capable?

Google DeepMind’s Levels of AGI framework treats breadth and depth of performance as distinct dimensions. Autonomy is a related deployment question.

What would persuade us?

We would look for broad, independently measured performance, with unfamiliar tasks and the comparison group specified.

Calling something an early level of AGI is a different claim from demonstrating expert performance across most domains.

Read Levels of AGI ↗

Showing Breadth of ability.

01 / What we can observe

The signals

Dated observations, the original sources and what each one actually establishes.

Read the current record ↗

02 / What else could happen

Alternative scenarios

Explore the optimistic, pragmatic and sceptical assumptions beside our central predictions.

Compare the possibilities ↗

03 / When our view changes

The review record

Original claims, revisions and withdrawals. A useful outlook has room to change its mind.

Inspect the decisions ↗
How we make the call · predictions, dates and uncertainty

These are our editorial predictions. We choose each target date by assessing current capability, the remaining obstacles and the time needed to demonstrate reliable results. The sources inform that judgement. They do not calculate a date or a number of days for us. Our initial predictions are expressed by year, using 31 December as their target. A later review can publish a different date when its reasoning supports that change.

Each prediction has its own published resolution criteria. Those thresholds are our definitions of a meaningful milestone, not universal industry standards. When evidence meets or fails them, we will record our assessment. A countdown reaching zero only means its deadline has passed; it never declares success automatically.

Each counter subtracts the current time from the end of its own published target date, in UTC. It shows full days remaining, or <1 day during the final 24 hours. It updates as days pass, when you return to the page and when a new review is published. Two predictions with the same target date have the same day count. A count in days does not make the underlying forecast precise to a day.

We first published these predictions on 15 September 2026. Earlier clocks showed illustrative scenario windows. Those remain in the alternative scenarios, with their original histories. They are not retroactively counted as our predictions. The earlier supporting forecasts also retain their own review and resolution criteria.

A changed prediction requires an appended, dated review. When we move a target, its day count is recalculated and its history shows how many days the review added or removed. That change compares the old and new target dates, so ordinary time passing cannot look like an evidence revision. Earlier dates, milestones, reasoning and evidence stay in the record. A changed milestone is labelled separately.

The scheduled evidence review reassesses our predictions each Monday. The open page checks for published reviews every five minutes and on return, retaining the last verified data if a check fails. Checking for a published update does not itself assess research or change a prediction. A site build does not renew a review date. See the review schedule →

Inspect the prediction history ↗ · Read the published prediction data ↗ · Explore the earlier turning points ↗