Watch enough AI video and you start seeing the same failure. The first shot is excellent. By the fourth shot the character has a slightly different nose, a different jaw, hair that behaves differently. Nobody can point at the exact frame where it happened, but everyone can feel that it did.

This is the single most common reason AI video reads as AI video. Fixing it is less about the model and more about how you set the job up.

CHARACTER 01 REFERENCE SHOT 1 SHOT 2 SHOT 3 SHOT 4 DRIFT
Audiences forgive small variance. They do not forgive a different person.

What does not work

Two things get recommended constantly and neither holds up on their own.

The first is fixed seeds. The intuition is reasonable: lock the randomness, get the same person. In practice, one of the clearer practitioner guides on this puts it plainly, that neither a repeated prompt nor a fixed seed guarantees continuity. A seed constrains one axis of variation in a system with many.

The second is copy pasting the same prompt. Necessary, but nowhere near sufficient. The same words produce a different person the moment the scene description around them changes.

If your consistency strategy is a seed and a saved prompt, you have a consistency hope, not a consistency strategy.

What actually works

Reference images are the load bearing element. Not the prompt. Every practical guide converges on this: a single clean portrait is enough for a short clip, and longer sequences need multiple angles. Two to four references is the working range for anything with several shots.

Frame chaining is the technique most people miss. Generate a clip. Export a clean frame from it. Use that frame as the reference for the next clip. Repeat. You are handing the model its own best output as the definition of the character, which stops identity drifting a little further with every generation. This single habit removes more drift than any prompt tuning.

Frame chaining

Each clip hands its best frame to the next, and the chain returns to the original references

Reference settwo to four referencesClip 1generateexport a clean frameClip 2that frame is the referenceexport a clean frameClip 3repeatthen go backto the original set
Handing the model its own best output as the definition of the character stops identity drifting a little further with every generation.

Split your prompt into a fixed block and a variable block. Write the character once, in a form you never touch:

Fixed: age, build, facial structure, hair, and the signature clothing or colour that identifies them. Variable: location, lighting, action, camera. Only the second block changes between shots.

The moment you start rewording the identity description to fit a new sentence, you have given the model permission to reinterpret the person. Treat the identity block as code you do not refactor.

The failure modes, and what each one means

  • The face changes gradually across shots. Your identity description is being reworded, or you are not chaining frames. Lock the block, chain the frames.
  • The face distorts during movement. The motion is too aggressive for the model to hold structure through. Reduce the intensity of the action or shorten the clip. Fast motion is where identity goes to die.
  • The character looks like a different person in a new location. Lighting is doing it. Keep colour temperature and key direction consistent between shots, or accept that you will be colour matching in post.
  • Everything drifts after the third or fourth clip. You are generating in a chain without re anchoring. Go back to your original reference set rather than referencing a reference of a reference.

Aim for perceptual continuity, not a pixel match

This is the reframe that makes the whole problem tractable. You are not trying to produce a pixel identical face across every frame. You are trying to produce a person an audience never questions.

Those are different standards, and the second one is achievable today. Real film has the same tolerance built in. Lighting changes, lenses change, makeup changes between setups shot weeks apart. Nobody notices, because the identity signals that matter stay fixed: facial structure, proportion, hair, and the visual signature the character carries.

Audiences forgive small variance. They do not forgive a different person.

Why this belongs in pre production

Everything above is cheaper before you generate anything. The reference set, the locked identity block, the lighting decision and the clothing signature are pre production work, and treating them as such is the difference between a shoot and a salvage operation.

Teams that build a character sheet first and generate second spend their time directing. Teams that generate first spend their time regenerating.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →