Instrument log · 1950 – present

AGI Watch

Seventy-six years of a field promising the same destination. This is not a release log — it carries only the moments that changed the answer to are we getting there: the capability thresholds, the setbacks that broke the last two schedules, and, increasingly, the claims that we have arrived. 20 entries in seventy-six years, so the ones that are here have to earn it.

70.2 years since Dartmouth named the field3.8 years since ChatGPT shipped9 days since OpenAI said “the AGI era” out loud

Threshold-moving events per half-decade — 1950 to now. Not release cadence: a quarter with four new models and no new answer counts zero

'50
'65
'80
'95
'10
'25

Where the frontier sits, on one framework

OpenAI's own five-level capability ladder, disclosed internally in July 2024. The marker is outside analysts' rough read, not an OpenAI score.

≈ here, per outside analysts
1Chatbots
2Reasoners
3Agents
4Innovators
5Organizations

Nobody agrees this ladder is the right one, or that “AGI” is a single threshold rather than a fuzzy region. Treat the marker as one dated opinion, not a reading off an instrument — the whole view is, honestly.

“The AGI Era”

2026–now

The industry starts saying the word out loud, on the record — while flagging its own model as a security risk.

Sep 3, 2026
AGI claim⚠ critical risk flag

GPT-6 “Astra”: “Welcome to the AGI era”

President Greg Brockman ends a press briefing with the line. Pressed on whether Astra actually is AGI, he calls the term a “gray, fuzzy thing,” says it is “not unreasonable to feel” we are now in that era, and leaves readers to decide. The gap between the sign-off and the answer is the entry. Astra is also the first model OpenAI has flagged at a “critical” cybersecurity risk threshold.

Sources: OpenAI, Axios, The New Stack

Feb 2026
AGI claim

Altman: "a couple of years away"

At the India AI Impact Summit, Sam Altman says the current trajectory may put us a couple of years from early versions of true superintelligence — and that systems have gone from struggling with high-school math to doing research-level mathematics and deriving novel results in theoretical physics. He also dates a threshold: by the end of 2028, more of the world's intellectual capacity inside data centres than outside.

Sources: MediaNama, Outlook Business

The Acceleration

2025–2025

Frontier capability stops being scarce. An open-weight model reaches it and is given away, which changes who the arrival question is even about.

Jan 2025
Breakthrough

DeepSeek R1 rattles markets

An open-weight reasoning model close to the frontier, released nearly free. A week later the market reads what it means: on 27 January, Nvidia falls about 17% and loses roughly $593bn of market value in a day — the largest single-day loss in Wall Street history at the time. Frontier capability turns out not to be scarce, which changes who the arrival question is even about.

Source: DeepSeek-AI, arXiv

Reasoning Arrives

2024–2024

Capability starts coming from time spent thinking, not only from training scale — and the first “levels of AGI” framework leaks.

Sep 2024
Breakthrough

o1-preview, the first "reasoner"

A model that visibly thinks step by step before answering, trading speed for deliberation. Capability starts coming from time spent at inference, not only from training scale — a second axis to grow along.

Source: OpenAI

Jul 2024
AGI claim

OpenAI's five levels of AGI, leaked

Bloomberg reports a framework OpenAI put to an all-hands, ranking progress from Chatbots through Reasoners, Agents, Innovators, to Organizations — the ladder this view borrows. Two things worth holding onto: the company grading the exam also wrote it, and at the time it told staff it was still on level one, on the cusp of two.

Source: Bloomberg

The Gold Rush

2023–2023

Every major lab ships a frontier model within months of the others, and exam scores become the currency of the claim. Those scores turn out to need reading carefully.

Mar 2023
AGI claim

"Top 10% on the bar exam"

OpenAI reports GPT-4 passing a simulated bar exam in the 90th percentile, and the figure travels everywhere. A later re-analysis finds that percentile was benchmarked against February repeat takers who had already failed once. Measured against July data it falls below the 69th percentile, to roughly the 62nd against first-time takers, and to roughly the mid-40s among people who actually passed — with essays weakest of all.

Sources: OpenAI, Martínez, Artificial Intelligence and Law

Language Models Go Mainstream

2020–2022

Generality starts falling out of scale rather than design — and then a chat window puts the question in everyone’s hands.

Nov 30, 2022
Release

ChatGPT launches

A chat interface on top of GPT-3.5. Nothing underneath it was new — the few-shot behaviour had been published two and a half years earlier — and that is the point: no threshold was crossed here, but the question stopped being a research one and became a public one. Adoption figures from this period are all third-party estimates, so this entry does not quote any.

Source: OpenAI

Nov 2020
Breakthrough

AlphaFold2 solves protein folding

Near-experimental accuracy on a 50-year-old grand challenge in biology — the strongest case on this page that a machine did novel science rather than fast lookup.

Source: DeepMind

May 2020
Release

GPT-3 learns tasks from the prompt

"Language Models are Few-Shot Learners": at 175B parameters it performs new tasks from a handful of examples in the prompt, with no retraining. The first serious evidence that generality might fall out of scale rather than design. The API followed in June.

Source: Brown et al., arXiv

The Deep Learning Turn

2012–2019

Learned representations beat hand-written ones, and the estimates for what is decades away start coming back badly wrong.

Feb 2019
AGI claim

GPT-2, held back as "too dangerous"

OpenAI stages the release rather than publishing outright, citing misuse risk. The first time a lab's own capability warning was itself the announcement — a genre this page now tracks closely.

Source: OpenAI

Mar 15, 2016
Breakthrough

AlphaGo beats Lee Sedol

DeepMind's system defeats a top Go champion at a game widely held to be decades from falling. The estimate was not wrong by a little.

Source: DeepMind

Sep 2012
Breakthrough

AlexNet wins ImageNet

A deep convolutional network trained on GPUs blows past every prior image-recognition result. Learned representations beat hand-written rules decisively, and the field changes method.

Source: Krizhevsky, Sutskever & Hinton (CACM reprint)

Quiet Machine Learning

1997–2011

No “AI” branding, just statistics and search beating humans one narrow game at a time.

Feb 16, 2011
Breakthrough

Watson wins Jeopardy!

IBM's question-answering system beats the show's two greatest champions on live television — the first time a machine handled open-domain language well enough to look general.

Source: IBM

May 11, 1997
Breakthrough

Deep Blue beats Kasparov

IBM's chess machine defeats the reigning world champion in a full match. Chess had been the standing proxy for machine intellect; passing it moved the goalposts instead of settling anything.

Source: IBM

Winters & Expert Systems

1974–1996

Two boom-bust cycles teach the field to be embarrassed about its own predictions.

1987–93
Setback

Second AI Winter

The market for specialized Lisp machines and expert systems collapses; "AI" becomes a word researchers avoid using. Two for two on the promises.

1980
Release

XCON goes into production

The first commercially successful expert system starts configuring computer orders for Digital Equipment Corp, saving millions — and restarts the argument that general intelligence is a matter of enough hand-written rules.

Source: McDermott, Artificial Intelligence 19(1)

1974–80
Setback

First AI Winter

Funding collapses after symbolic AI fails to deliver on its promises; the Lighthill report guts UK research money. The field's confidence had been the loudest evidence for its timelines.

Source: The Lighthill Report, 1973

Foundations

1950–1973

The question gets asked, the field gets a name, and the first schedule is set.

Jan 1966
Release

ELIZA, the first chatbot

Joseph Weizenbaum's pattern-matching therapist script convinces some users they're talking to something that understands them. The gap between seeming and being opens here, and never closes.

Source: Weizenbaum, CACM 9(1)

Summer 1956
Breakthrough

Dartmouth Workshop names the field

McCarthy, Minsky, Shannon and colleagues coin "artificial intelligence" — and propose that ten people could make significant progress on language and abstraction in a single summer. The first schedule, and the first one missed.

Source: The 1955 proposal

Oct 1950
Breakthrough

Turing asks "can machines think?"

The Imitation Game paper proposes a behavioral test for machine intelligence, before the field even has a name. Every argument on this page is downstream of it.

Source: Turing, Mind LIX(236)