On September 28, Kling, the AI video company spun out of the Chinese platform Kuaishou, announced Kling 4.0. A faster “Flash” version is live now for its top annual subscribers, and the full model is promised for October. The announcement reads like a wish list for anyone who makes video with AI.
What was announced
On picture and sound, Kling promises more stable motion, stereo audio, more accurate lip sync, output up to 4K and 10-bit HDR. On control, an “Omni Reference” system accepts up to 15 references across images, video and audio, with “more consistent results” as the stated goal, and up to 10 keyframes can steer a single clip. The full model generates 30 seconds natively; the Flash version, built for speed and cost, stops at 20. Video extension and support for several languages, accents and dialects round out the list.
All of this is the vendor’s word. At the time of writing, no independent test has measured any of it.
How long one generation can run
Maximum clip length, as announced
| Version | Maximum length |
|---|---|
| Kling 4.0 Flash | 20 s |
| Kling 4.0 | 30 s |
The money behind the race
According to Bloomberg, Kling raised $2.8 billion in July from Alibaba, Baidu and Tencent, for a valuation of about $18 billion, and its annualized revenue passed $500 million in March. The target is ByteDance, whose Seedance 2.5, released in July, has become a default for many creators, with Alibaba and MiniMax also in the race.
For a business, the lesson is not which model wins this month. It is that a new leader shows up every few months. The best tool today will not be the best tool next spring. We saw what that means when a major model shut down.
What really changes for a brand
Consistency, if the references hold. Fifteen references are aimed squarely at the first failure of AI video: a product, a face or a set that drifts from one shot to the next. If it works as announced, fewer shots end up in the bin. Why that problem is so stubborn.
Room to grade. For anyone who color-grades, 10-bit HDR matters more than 4K. It keeps detail in skies and shadows when the image is pushed, which brings generated footage closer to what broadcast and cinema pipelines expect.
Cheaper language versions. Voices, lip sync and dialects in one model make localized versions cheaper to produce. For Quebec, the real test is whether “dialects” includes a credible Québécois French. Until someone tests it, assume it does not.
What does not change
A 30-second shot is not a 30-second ad. A good spot is built from several shots, and its power sits in the cuts between them: what the viewer sees first, what comes next, what it all means. A 100-year-old experiment explains why. A model that generates 30 continuous seconds generates a long take, and a long take needs a reason to keep rolling. Even the famous one-shot films cheat.
The ten keyframes say the same thing in their own way. They exist because someone still has to decide what happens at each moment. The model asks for direction more explicitly. It does not supply it.
What to do with it
If you buy video, do not pick a studio for the model it uses: models change every few months. Ask to see a result on a product like yours, and ask how they keep your product and your characters consistent from one shot to the next. That answer outlasts any version number.
If you make video, test Flash only if you already have access, and wait for independent tests of the full model before rebuilding anything around it. The promises are real progress. The proof comes in October.
Sources
- Kling 4.0 announcement (Kling AI on X, September 28, 2026)
- Kling AI Launches Kling 4.0 with 30 Second Video Generations (HuggingNews, September 28, 2026)
- Kling 4.0 Debuts 30s AI Video Model, reporting Bloomberg (Briefs, September 28, 2026)
- Kuaishou’s AI Video Spinoff Unveils New Model in ByteDance Chase (Bloomberg, September 28, 2026)