Ask five people working in AI whether we have AGI and you will get five answers, all confident. That is not because some of them are uninformed. It is because they are answering different questions.
In April 2026, Helen Toner published the sharpest statement of the problem: the term is "almost useless at this point." Her argument is that capability now sits in what she calls a "fuzzy cloud of AGI-ish" that satisfies some definitions and fails others, so serious people can simultaneously claim AGI has arrived and that it is a decade out, without either of them being wrong on their own terms. Her line is worth keeping: "Expecting that we'll know it when we see it is patently not working."
Six definitions, all in active use
These are not strawmen. Each is load-bearing for someone.
- OpenAI's Charter: "highly autonomous systems that outperform humans at most economically valuable work."
- Google DeepMind's Levels of AGI: not a threshold at all, but a matrix of performance against generality, with six tiers. On that framework, ChatGPT and Gemini already sit at Level 1, "Emerging AGI."
- Hendrycks et al.: matching "the cognitive versatility and proficiency of a well-educated adult," measured across ten weighted cognitive domains.
- François Chollet and ARC Prize: matching the learning efficiency of humans, not their scores. Intelligence as skill acquisition, not skill.
- Anthropic: avoids the word. Dario Amodei writes about "powerful AI" instead, defined as smarter than a Nobel laureate across most relevant fields.
- The Forecasting Research Institute: an operational definition for surveys, outperforming 90th-percentile professionals on 90% of non-physical tasks at no more than five times human labour cost.
Those are not variations on a theme. DeepMind's says we are already inside it. Chollet's says scores are the wrong measurement entirely. Anthropic's refuses the frame.
Change the maths, move the answer thirty points
The Hendrycks framework is the most cited quantified definition of the last two years, signed by 33 researchers including Yoshua Bengio, Gary Marcus, Max Tegmark and Erik Brynjolfsson. It scores GPT-4 at 27% and GPT-5 at 57% across ten cognitive domains.
Then Fares Fourati at KAUST pointed out that averaging those ten domains arithmetically assumes they compensate for each other, that being excellent at recall makes up for being unable to retain anything. Recompute the same scores without that assumption and GPT-4 falls from 27% to 7%, GPT-5 from 58% to 24%.
Nothing about the models changed. The definitional choice moved the number by more than thirty points.
That is the whole problem in one example. A number that swings that far on a modelling assumption is not measuring a property of the system. It is measuring an argument.
What the systems actually cannot do
Underneath the definitional fight there are concrete, measured gaps, and they are more useful than the label. The Hendrycks work found a "highly jagged" profile: strong on knowledge, with critical deficits in what the authors call foundational cognitive machinery. Long-term memory storage is the clearest one, where both GPT-4 and GPT-5 scored zero.
The jaggedness shows up in embarrassing places. The 2026 Stanford AI Index recorded GPT-5.4 scoring 50.6% on ClockBench, which asks models to read analogue clocks, while Claude Opus 4.6 scored 8.9%. These are systems that sit near the ceiling on graduate-level physics.
And on genuine novelty the gap is stark. When ARC-AGI-3 launched in March 2026, testing whether a system can work out the rules of an unfamiliar environment by exploring it, frontier AI scored 0.51% against a 100% human baseline.
The definition stopped being worth money
Here is the detail that says most about where the term has landed. AGI used to have a contractual meaning. The Microsoft-OpenAI agreement contained a clause tied to it, reportedly pegged at one point to $100bn in cumulative profits, later to verification by an independent expert panel.
On 27 April 2026 that arrangement was replaced. Microsoft's licence became non-exclusive and revenue share now runs through 2030 "independent of OpenAI's technology progress." The single place where AGI had a legally consequential definition quietly stopped needing one.
The question that actually helps
Arvind Narayanan and Sayash Kapoor make the argument that matters for anyone running a business: "If a company declares that it has built AGI, based on whatever definition, it is not an actionable event." No observable threshold, no immediate implications, and decades of diffusion between capability and effect.
So the useful question is not whether a system qualifies. It is narrower and answerable: on the specific work you do, what can this thing complete reliably, at what cost, and how would you know when it was wrong? That question has an answer this quarter. The other one has had six answers for three years.
Sources
- Helen Toner, "The term 'AGI' is almost useless at this point" (6 April 2026)
- Morris et al., "Levels of AGI: Operationalizing Progress on the Path to AGI" (Google DeepMind)
- Hendrycks et al., "A Definition of AGI"
- Fourati, "A Coherence-Based Measure of AGI"
- ARC Prize, ARC-AGI-3 launch (25 March 2026)
- IEEE Spectrum on the Stanford AI Index 2026 (13 April 2026)
- OpenAI, "The next phase of our Microsoft partnership" (27 April 2026)
- Kapoor & Narayanan, "AGI is not a milestone"