The common framing is a ladder: narrow AI today, general AI later, with a line between them we will cross on some announced date. The frameworks built by the people who work on this do not describe it that way at all.

Generality is a dimension, not a door

Google DeepMind's Levels of AGI plots two things against each other: performance and generality. Narrow systems and general systems both appear, at six levels of capability each. The consequence is immediate and counterintuitive. AlphaFold and Stockfish sit at Level 5, Superhuman, in the narrow column. ChatGPT and Gemini sit at Level 1, Emerging, in the general column.

So a system can be radically superhuman and radically narrow simultaneously. Chess engines passed every human two decades ago and are no closer to general intelligence than they were. That is why "is it AGI yet" resists a yes or no: you are asking for a single coordinate from a system that has two.

The frontier is jagged, not a front line

Ethan Mollick's term for what this feels like in practice is the jagged frontier. A model can be superhuman at differential medical diagnosis and fail a simple visual puzzle, and there is no principled way to predict which from the outside.

The 2026 Stanford AI Index gives the cleanest illustration. On ClockBench, which asks a model to read an analogue clock, GPT-5.4 scored 50.6% and Claude Opus 4.6 scored 8.9%. Both models are near the ceiling on benchmarks designed to test graduate-level reasoning. A five-year-old does the clock task.

"AI can do X" is never a general statement. It is a claim about one task, one model, one week.

Where the gaps sit

The Hendrycks framework measured ten cognitive domains and found what the authors call a highly jagged profile. The systems are strong on general knowledge and reading. They are weak or absent on what the paper describes as foundational machinery, above all long-term memory storage, where both GPT-4 and GPT-5 scored zero.

That single gap explains a lot of day-to-day frustration. A model that cannot consolidate what it learned yesterday is not a colleague who is getting better at your account. It is a very well-read stranger, every morning.

Why this matters commercially

Because it kills the two most common planning errors at once.

  • Extrapolating from a win. A model that drafts your briefs beautifully tells you nothing about whether it can reconcile your invoices. Different point on the frontier.
  • Extrapolating from a failure. A model that cannot read your analogue dashboard is not therefore unreliable at summarising a contract.

Measured exposure and measured adoption come apart for exactly this reason. The St. Louis Fed found that medical secretaries adopt generative AI at 16.8%, against 61% predicted by task exposure alone, because privacy and the cost of an error dominate the theoretical fit. What a system could do and what it is trusted to do are separate measurements.

The practical read

Stop asking where a model sits on the ladder. Map the jaggedness on your own work instead: take ten tasks you actually do, run them, and record what came back reliably, what came back plausibly but wrong, and what cost more to check than to do. That map is worth more than any capability announcement, and it is the only version of the question that changes what you do on Monday.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →