On a set you light the space once, then shoot the wide, the mid, the singles and the insert inside that setup. You cheat lamps between setups but you never break the logic: the window is still camera-left in the reverse.

AI video has no equivalent of that. Every generation is an independent sample. The model is not holding your scene in memory between calls, it is re-deriving a plausible lighting state from your prompt each time. Anything you did not pin gets re-rolled, and lighting is high-dimensional and usually unspecified, so it drifts hard.

Lightricks says so in its own documentation, about cuts inside a single generation: "Restate scale, angle, framing, and any lighting change at the cut. The model does not inherit the previous shot's setup." Across separate generations it is strictly worse.

SC 04 · COVERAGE SHOT A SHOT B CUT KEY · CAM LEFT KEY · CAM RIGHT
Every generation re-derives the lighting: the model does not inherit the previous shot’s setup.

The craft vocabulary is the tool, not the lamps

None of the standard cinematography terms stop being useful. What changes is their function. You are not instructing a gaffer, you are building a controlled vocabulary you can restate identically across prompts.

  • Key, fill, back. Not lamps you rig, but three roles precise enough to re-serve verbatim.
  • Motivation. Every light explicable by a source in the scene. This is the most transferable idea, because motivation is a fact about the world you can write down once. "Moody" is a fact about your feelings and cannot be repeated.
  • Colour temperature in kelvin. Tungsten practicals around 2700 to 3200K, daylight around 5600K. Mixed temperature is what makes an interior read as real.
  • Direction, camera-relative. Camera left, camera right. Camera-relative is what survives being written into a prompt.
  • Quality. Hard or soft, governed by source size relative to subject.
  • Contrast ratio. Naming 3:1 or 8:1 is more repeatable than naming a mood.

Notice that when the labs write their own good examples, this is exactly what they write. OpenAI's Sora 2 guide contrasts a vague "brightly lit" against "soft window light with warm lamp fill and a cool edge from the hallway". That is key, fill and edge, in plain English. Google's Veo examples name "harsh fluorescent overhead lights and the green glow of the monochrome monitor", which is a mixed-practical interior described by source and colour.

Every lab's best example names a source, a direction and a colour. None of them names an adjective.

Where lighting sits in a prompt

It is now a first-class field in several published prompt structures. Sora 2's template carries a dedicated line: Lighting + palette: [specific quality, direction, color sources]. Seedance 2.5's structure has LIGHTING as its own labelled block, specified as motivated source, direction and surface behaviour. Kling's published formula names lighting explicitly, though its documented vocabulary is thin and mood-level.

One notable exception, and it cuts against common practice. Runway's Gen-4 video guide contains no lighting vocabulary at all, on the stated basis that the input image already carries colour and lighting and the text should describe motion. It goes further: restating elements already in the image "can lead to reduced motion or unexpected results." In Runway image-to-video, repeating your lighting block is not redundant, it may actively cost you motion quality. There is no single universal rule here.

What actually holds continuity

In descending order of how well documented it is.

  • Extension, conditioned on real footage. The strongest mechanism available, because the model is looking at frames rather than re-imagining them. Veo 3.1's extend conditions on the last second of video. The limit is fundamental: extension buys continuity along a shot, not across a cut. It does not solve coverage.
  • First and last frame conditioning. Documented, narrow, effective for transitions.
  • Reference images. Clearly help in practice. But be precise about what is documented: Veo's ingredients, Runway's references, Kling's elements and Seedance's reference roles are all documented for subjects, characters, objects and style. Runway's own copy advertises holding a character consistent across different lighting conditions, which is the opposite of holding lighting. No vendor documents lighting persistence between separately generated clips as a supported behaviour. You are riding a style mechanism to do a lighting job.
  • Repeating an identical lighting block verbatim. No model exposes a lighting lock, so this is practitioner practice rather than a feature. But it is practice the vendors' own prompt architecture is built to accommodate, which is a fair way to put it.

Seeds do not do what people think

The persistent folklore is that holding a seed constant across different prompts keeps lighting matched between shots. Nothing supports this. A seed indexes the noise sample, not the scene. Change the prompt and you change the trajectory.

Worse, seeds are not even reliably deterministic on hosted models. Runway's own API reference says an identical request with the same seed "will produce similar results." Similar, not identical. And no vendor promises a seed survives a model update. Seeds are good for A/B-ing one prompt against itself. They are not a continuity tool.

The reverse-angle trap

This one is reasoning rather than a sourced finding, and it is the thing most likely to catch someone who already knows cinematography.

Your lighting block is camera-relative. On a real set, a reverse shot flips screen direction: a key from camera left in the over-the-shoulder becomes a key from camera right in the reverse, because the lamp did not move, the camera did. Paste your block verbatim into the reverse and the key stays on camera left, silently breaking the geography of the room.

Two ways out. Hand-flip the direction on every reverse, or write the block in world-relative terms, the window is camera-left in the wide and behind camera in the reverse, and accept that you are maintaining a small lighting continuity sheet by hand. Either way it is deliberate work, and skipping it is where most sequences fall apart.

Time of day, and a trick worth stealing

Bind the light to a fact rather than a look. The Seedance guidance puts the reason well: golden hour keeps the light the same across every segment without describing the sun's position in each one. A named time of day is a compact constraint the model has enormous prior for. A described sun elevation is a long instruction it may quietly ignore.

For practicals, exploit what is documented. Reference and element features lock objects far better than they lock illumination. So describe a practical lamp as a physical object with a fixed position and state, "the brass desk lamp at the right of the desk, shade tilted toward the wall, switched on, warm 2700K," and you recruit the object-consistency machinery to do lighting work.

Where it breaks hardest

Exterior day with a single dominant source is easiest: one named condition constrains nearly everything. Interiors are harder because they are multi-source, and multi-source is where the documented failures cluster. Night is hardest, being low-signal and dominated by practicals.

There is real published evidence here, though it needs care. PhyWorldBench, a 2026 benchmark with 12 models and 12,600 human-annotated videos, includes "Lighting and Shadows" as one of its ten physical categories, which is itself a citable fact about where the field thinks the problems are. Its named failures include a flashlight moving across a dark room that does not illuminate its surroundings, and light failing to refract. Only about a quarter of videos satisfied both semantic and physical correctness. Physics-IQ reaches a blunter conclusion: visual realism does not imply physical understanding.

Neither benchmark publishes an isolated lighting score, so nobody should quote one.

But there is a counterweight that reframes the whole problem. A 2026 paper from Peking University, BAAI and OpenBayes argues that video diffusion models implicitly encode complex physics, geometry and material optics, well enough that the authors extract usable HDR environment maps from them. If that holds, lighting continuity is not a knowledge failure in these models. It is a control failure. They know more about light than you can currently address through a prompt.

Generate the look, grade the balance

The practical split, and the constraint that decides it. Lighting direction, quality, contrast and motivation are baked into pixels and cannot be recovered in a grade. No amount of colour correction moves a key light. Those must be right in generation.

Colour temperature, overall balance and shot-to-shot trim are grade work, and chasing them in-prompt costs re-rolls for something a colourist fixes in minutes.

How forgiving that second step is depends on bit depth, and this is the number to know: most AI video tools output 8-bit by default, 256 levels per channel against 1024 at 10-bit, which means banding and posterisation the moment anyone lifts the blacks. Sixteen-bit scene-linear EXR output is now appearing, and where it is available you can generate looser and grade harder. Where it is not, more has to be right in camera, so to speak.

Most AI video tools output 8-bit by default

Levels per channel, 8-bit vs 10-bit

8-bit: 256 levels8-bit: 256 levels8-bit256levels 10-bit: 1024 levels10-bit: 1024 levels10-bit1024levels
Most AI video tools output 8-bit by default
Bit depthLevels per channel
8-bit256
10-bit1024
Source: Lightricks, AI video for broadcast and the 8-bit constraint (12 March 2026).

A middle layer also now exists that did not before. Video-to-video relighting can change the time of day or soften a harsh key after generation. It is a re-generation with its own drift risk, so the operational rule from the documentation applies: pin everything else in place explicitly, or it will move.

Failure modes, named

Only the first of these is an established term of art. The rest are a working vocabulary, offered as one, not claimed as industry standard.

  • Flicker. Frame-to-frame instability in intensity and colour within one clip. Worst at night and in high contrast. Scrub slowly to catch it.
  • Relight-on-cut. Two shots that should share a setup arrive lit differently. The core failure, and directly attributable to the model not inheriting the previous setup.
  • Shadow direction flip. The most visible form of the above, and the failure your own verbatim prompt block causes on a reverse angle.
  • Practical state change. A lamp on in one shot and off in the next, or moved, or a different colour. Models treat practicals as set dressing rather than continuity-critical props.
  • Exposure pumping. Overall brightness drifting within a shot as the subject moves, like an auto-exposure hunting. Low-frequency, follows the action, distinct from flicker.
  • Missing contact shadow. No shadow where the subject meets the floor. The fastest tell that a shot is generated, and the reason a composite reads as pasted together even when nothing else is wrong.
  • Unmotivated source. Light arriving from nowhere the scene accounts for. The direct violation of the motivation principle, and the reason to write motivation into the prompt as a fact.
  • Source that fails to illuminate. The inverse: a light is rendered but does not affect its surroundings. The characteristic night-interior failure, and documented in the benchmark literature.

The short version

Write the lighting down once, as a block naming source, direction, quality, colour temperature and contrast ratio. Restate it verbatim on every shot except Runway image-to-video, and hand-flip the direction on every reverse. Use extension where you can, because footage-conditioned continuity beats prompt-described continuity every time. Get direction and motivation right in generation because the grade cannot fix them, and leave balance to the grade. Then watch for the eight failures above, because knowing their names is most of catching them.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →