Read a hundred AI video prompts and you will notice almost all of them describe a subject. A woman in a red coat walking through a train station. A watch on a marble surface. A car on a coast road.
None of that is a shot. It is a description of something that exists. The camera, which is the thing that turns an existing thing into a piece of film, is missing entirely.
The good news is that these models already speak the language. You just have to use it.
The vocabulary they understand
Camera terminology from actual filmmaking is not decoration in a prompt. It is the most reliable control surface you have, because it appears throughout the material these models learned from. Runway's own guidance lists the terms that work, and they are exactly the ones a camera operator would use:
- Movement: dolly in and push in, dolly out and pull back, pan left and right, whip pan, tilt up and down, truck left and right, pedestal up and down, orbit, arc, crane up, drone rise, crash zoom, handheld, steadicam, gimbal, and locked off static.
- Angle: eye level, low angle, high angle, over the shoulder, POV, top down and aerial, bird's eye, worm's eye, and Dutch angle.
- Shot size: extreme close up and macro, close up, medium, full shot, wide and establishing, extreme wide.
Every one of those terms carries meaning a vague description cannot. "Low angle" tells the model where the horizon sits, how the subject reads in terms of power, and what the lens sees of the ceiling or sky. "Cinematic" tells it nothing.
The order that works
Structure matters as much as vocabulary. The ordering Runway recommends is worth memorising:
Shot size, then angle, then movement with direction and speed, then subject and action, then lens and look, then lighting and mood, then what the shot reveals.
Applied to the train station example, the difference is stark. Instead of a woman in a red coat walking through a train station, you get: medium shot, eye level, slow dolly push in over four seconds, a woman in a red coat walking toward camera through a crowded station concourse, 35mm, hard morning light from high windows, revealing that she is the only person not moving with the crowd.
That second version is a directing instruction. It has an intent. It also tells you whether the model failed, because you specified what success looks like.
The order that works
Runway’s recommended prompt order, applied to the train station example
| Step | Element | Example |
|---|---|---|
| 1 | Shot size | medium shot |
| 2 | Angle | eye level |
| 3 | Movement + speed | slow dolly push in over four seconds |
| 4 | Subject and action | a woman in a red coat walking toward camera |
| 5 | Lens and look | 35mm |
| 6 | Lighting and mood | hard morning light from high windows |
| 7 | Reveal | the only person not moving with the crowd |
Always give movement a speed
This is the highest leverage detail in the whole guide, and the one most people skip. "Push in" is an instruction with no rate attached, so the model picks one, usually badly. "Slow dolly push in over four seconds" constrains it.
The same applies to the relationship between camera and subject. Verbs like "follows" or "matches pace" tell the model that the camera is tracking something rather than moving independently near it. Untethered camera movement is one of the clearest tells of generated video.
One move per shot
The instinct is to ask for everything: a crane down into an orbit that ends in a push in. It almost never survives generation intact.
The better approach for a complex move is to describe what gets revealed rather than stacking commands. A camera that cranes down to reveal a figure in a doorway is one idea, and the model can hold one idea. A camera performing three choreographed moves in eight seconds is three ideas competing for the same frames.
If you genuinely need the complex move, break it into separate shots and cut them together. That is what an editor would do with real coverage anyway.
This is a directing skill, not a prompting trick
Worth being blunt about why this matters commercially. Anyone can generate a clip. The thing that separates professional output from the flood of generic AI footage is coverage and intent: shots that were chosen for a reason, that cut together, that put the camera where the story needs it.
The vocabulary above is the same vocabulary a director of photography uses on set, which is the point. The tool changed. The craft did not. Teams that already knew how to describe a shot have an enormous head start here, and teams that never learned are producing footage that technically works and communicates nothing.
Learn the grammar. It outlasts every model on the market.