Generative AI produces content: text, images, video, audio, code. That is a real and commercially enormous capability, and it is the one almost everyone is actually buying. The question worth asking is whether getting better at producing output moves a system toward general intelligence, or whether those are different directions.
The honest answer from the measurement record is that they are largely different directions, and the evidence is specific rather than philosophical.
Four gaps that generation does not close
Each of these is measured, not asserted.
- It does not learn from yesterday. On the ten-domain cognitive framework published by Hendrycks and colleagues, long-term memory storage is the one broad domain where both GPT-4 and GPT-5 score zero. The model that helped you on Tuesday starts Wednesday knowing nothing about Tuesday.
- It does not know when to stop. GPT-5 hallucinates in response to more than 30% of questions on SimpleQA. Calibrated abstention, the ability to say "I don't know," is a different skill from fluent production, and generation does not supply it.
- It does not model the physical world. On IntPhys 2, which tests intuitive physics, the best current models perform only slightly better than chance.
- It does not adapt to genuine novelty. ARC-AGI-3 puts a system in an unfamiliar environment and asks it to work out the rules by exploring. At launch in March 2026, frontier AI scored 0.51%. The human baseline is 100%.
None of these are edge cases. They are the difference between a tool that produces and a mind that operates.
Why fluency reads as understanding
Generative systems are optimised to produce output that looks like what a competent person would produce. When that works, the natural inference is that something competent produced it. That inference is exactly what the benchmark record does not support, and it is why the failures are so disorienting when they arrive: the register of the output does not degrade when the substance does.
The output stays confident while the correctness collapses. That is a property of generation, not a bug in a particular model.
The contested version
It would be dishonest to present this as settled. Sam Altman, discussing GPT-6 in August 2026, said he expects it to be "the first model where the model actually invents new things in a way that matters. That's a very AGI-like thing." That is a forward-looking claim from an interested party rather than a published result, and it should be read as such. But the position that scale in generation eventually produces discovery is a real position held by serious people, and the next two years will test it.
What this means if you are building on these tools
It changes what you put around the model rather than which model you pick.
- Memory is your job, not the model's. If continuity matters, it has to live in your systems: specs, reference sheets, locked briefs. Nothing carries over on its own.
- Verification is your job too. A system that cannot reliably abstain needs a step that checks, and that step has a cost you should be measuring.
- Novelty is where it will fail hardest. The further a task sits from anything in the training distribution, the worse the economics get, and the failure will not announce itself.
That is not a case against generative tools. It is a case for building the surrounding system deliberately, because the parts that generation does not supply are exactly the parts that decide whether the output is usable.