Read a hundred AI video prompts and you will notice almost all of them describe a subject. A woman in a red coat walking through a train station. A watch on a marble surface. A car on a coast road.

None of that is a shot. It is a description of something that exists. The camera, which is the thing that turns an existing thing into a piece of film, is missing entirely.

The good news is that these models already speak the language. You just have to use it.

MEDIUM · EYE LEVEL · 35MM CAMERA SUBJECT SLOW DOLLY PUSH IN · 4 SEC
Camera terminology from actual filmmaking is not decoration in a prompt. It is the most reliable control surface you have.

The vocabulary they understand

Camera terminology from actual filmmaking is not decoration in a prompt. It is the most reliable control surface you have, because it appears throughout the material these models learned from. Runway's own guidance lists the terms that work, and they are exactly the ones a camera operator would use:

  • Movement: dolly in and push in, dolly out and pull back, pan left and right, whip pan, tilt up and down, truck left and right, pedestal up and down, orbit, arc, crane up, drone rise, crash zoom, handheld, steadicam, gimbal, and locked off static.
  • Angle: eye level, low angle, high angle, over the shoulder, POV, top down and aerial, bird's eye, worm's eye, and Dutch angle.
  • Shot size: extreme close up and macro, close up, medium, full shot, wide and establishing, extreme wide.

Every one of those terms carries meaning a vague description cannot. "Low angle" tells the model where the horizon sits, how the subject reads in terms of power, and what the lens sees of the ceiling or sky. "Cinematic" tells it nothing.

The order that works

Structure matters as much as vocabulary. The ordering Runway recommends is worth memorising:

Shot size, then angle, then movement with direction and speed, then subject and action, then lens and look, then lighting and mood, then what the shot reveals.

Applied to the train station example, the difference is stark. Instead of a woman in a red coat walking through a train station, you get: medium shot, eye level, slow dolly push in over four seconds, a woman in a red coat walking toward camera through a crowded station concourse, 35mm, hard morning light from high windows, revealing that she is the only person not moving with the crowd.

That second version is a directing instruction. It has an intent. It also tells you whether the model failed, because you specified what success looks like.

The order that works

Runway’s recommended prompt order, applied to the train station example

1. Shot size: medium shot11. Shot size: medium shotShot sizemedium shot2. Angle: eye level22. Angle: eye levelAngleeye level3. Movement + speed: slow dolly push in over four seconds33. Movement + speed: slow dolly push in over four secondsMovement + speedslow dolly push in over four seconds4. Subject and action: a woman in a red coat walking toward camera44. Subject and action: a woman in a red coat walking toward cameraSubject and actiona woman in a red coat walking toward camera5. Lens and look: 35mm55. Lens and look: 35mmLens and look35mm6. Lighting and mood: hard morning light from high windows66. Lighting and mood: hard morning light from high windowsLighting and moodhard morning light from high windows7. Reveal: the only person not moving with the crowd77. Reveal: the only person not moving with the crowdRevealthe only person not moving with the crowd
The order that works
StepElementExample
1Shot sizemedium shot
2Angleeye level
3Movement + speedslow dolly push in over four seconds
4Subject and actiona woman in a red coat walking toward camera
5Lens and look35mm
6Lighting and moodhard morning light from high windows
7Revealthe only person not moving with the crowd
Source: Runway: AI Camera Prompts, How to Control Camera Movement, Angles and Shots.

Always give movement a speed

This is the highest leverage detail in the whole guide, and the one most people skip. "Push in" is an instruction with no rate attached, so the model picks one, usually badly. "Slow dolly push in over four seconds" constrains it.

The same applies to the relationship between camera and subject. Verbs like "follows" or "matches pace" tell the model that the camera is tracking something rather than moving independently near it. Untethered camera movement is one of the clearest tells of generated video.

One move per shot

The instinct is to ask for everything: a crane down into an orbit that ends in a push in. It almost never survives generation intact.

The better approach for a complex move is to describe what gets revealed rather than stacking commands. A camera that cranes down to reveal a figure in a doorway is one idea, and the model can hold one idea. A camera performing three choreographed moves in eight seconds is three ideas competing for the same frames.

If you genuinely need the complex move, break it into separate shots and cut them together. That is what an editor would do with real coverage anyway.

This is a directing skill, not a prompting trick

Worth being blunt about why this matters commercially. Anyone can generate a clip. The thing that separates professional output from the flood of generic AI footage is coverage and intent: shots that were chosen for a reason, that cut together, that put the camera where the story needs it.

The vocabulary above is the same vocabulary a director of photography uses on set, which is the point. The tool changed. The craft did not. Teams that already knew how to describe a shot have an enormous head start here, and teams that never learned are producing footage that technically works and communicates nothing.

Learn the grammar. It outlasts every model on the market.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →