Agentic AI is routinely described as the next rung up from generative AI, one step closer to general intelligence. The framework that introduced the levels of AGI describes it as something else entirely: a separate axis.
Two tables, not one ladder
DeepMind's Levels of AGI paper publishes capability and autonomy as two distinct scales. The capability scale runs from Emerging to Superhuman. The autonomy scale runs separately: AI as Tool, as Consultant, as Collaborator, as Expert, and finally as Agent, fully autonomous, at Level 5.
The paper's position is that higher capability unlocks higher autonomy without determining it. A highly capable system can deliberately be deployed as a tool for safety reasons. A limited system can be handed far more autonomy than it deserves. Risk is indexed to the pair, not to capability alone.
On the framework's own terms, "agentic" is an interaction-design decision. Somebody chose how much rope to give it.
The OECD's February 2026 landscape paper builds the same distinction into something a buyer can use: four levels of action autonomy, from a system that only recommends, to one requiring approval before acting, to one that acts unless stopped, to one that is fully autonomous. It also separates autonomy, acting without human involvement, from agency, forming goals and reasoning over long horizons. Most products sold as agentic have the first and not the second.
The measurement that settles the argument
In September 2026, ARC Prize published results for GPT-6 Astra on ARC-AGI-3, their benchmark for exploring unfamiliar environments. The model scored 62.7% under ARC Prize's standard harness. The same model, the same benchmark, the same day, scored 99.9% under a different harness.
A 37-point swing with no change to the model. Whatever that measures, a large part of it is the scaffolding rather than the intelligence. ARC Prize said so themselves, and their caution is worth quoting given it concerns their own best-ever result: "we are not claiming that it is AGI," and "saturating the benchmark would not represent proof of achieving AGI."
The case that agency is a route to generality
There is a real argument on the other side, and ARC Prize is making it with their benchmark design. ARC-AGI-3 defines four components of agentic intelligence: exploration, modelling, goal-setting, and planning and execution. It measures action efficiency, how many moves a system needs to solve an environment it has never seen. On that measure Astra used 51.7% fewer actions per level than the human baseline.
The claim behind that design is that following instructions and genuinely exploring are different things, and that the gap between them is where generality lives. That is not a trivial position.
The case that it is scaffolding
Gary Marcus and co-authors put the counter-argument in February 2026: current systems "lack persistent goals, struggle with long-horizon reasoning, and depend extensively on human scaffolding for task formulation, evaluation, and correction." On that reading, agentic performance is substantially borrowed from the harness, which is precisely what the 62.7% to 99.9% swing demonstrates.
Both can be true. Exploration under novelty may be necessary for generality and nowhere near sufficient, while most of what is currently shipped as agentic is scaffolding doing the heavy lifting.
What to take from this
When a vendor says agentic, they are describing a permission level, not a capability tier. The question that gets you somewhere is the one the OECD ladder forces: what happens when this is wrong, and who has to approve it before it acts? That is answerable, it is the actual risk, and it has nothing to do with where anyone thinks we are on the road to AGI.
Sources
- Morris et al., "Levels of AGI" (Google DeepMind), including the Levels of Autonomy table
- OECD, "What agentic AI is and does" (March 2026), companion to OECD AI Papers No. 56
- ARC Prize, "Astra" results on ARC-AGI-3 (3 September 2026)
- ARC Prize, ARC-AGI-3 Technical Report (22 April 2026)
- Gary Marcus et al., "Rumors of AGI's arrival have been greatly exaggerated" (17 February 2026)