In the late 1910s and early 1920s, a young Soviet filmmaker named Lev Kuleshov ran one of the simplest experiments in film history. It still explains more about how films work than most courses do, and it has never been more relevant than now, when anyone can generate a shot.

REC SOUP COFFIN DIVAN SAME FACE
Three different shots. One face that never changes.

One face, three emotions

Kuleshov took a shot of the actor Ivan Mosjoukine looking straight ahead with no expression at all. He cut it after three different shots: a bowl of soup, a girl in a coffin, a woman lying on a divan. Audiences saw hunger, then grief, then desire.

According to the filmmaker Vsevolod Pudovkin, who later recounted the experiment, viewers raved about the acting and were moved by the sorrow in the actor’s eyes. The face was exactly the same every time. Kuleshov’s conclusion: cinema is made of fragments, and of the way those fragments are put together.

Hitchcock’s version

In 1964, in a Canadian television interview, Alfred Hitchcock explained the idea with his own face. A shot of him squinting, then smiling. Cut it after a mother holding her baby, and he is a kind old man. Replace the mother with a young woman in a bikini: same face, same smile, and now he is a dirty old man. Nothing in his shot changed.

Same face, different meaning

What audiences read into a shot, depending on the shot before it

Bowl of soup → hungerBowl of soup + the face→ hungerGirl in a coffin → griefGirl in a coffin + the face→ griefWoman on a divan → desireWoman on a divan + the face→ desireMother and baby → a kind old manMother and baby + the face→ a kind old manWoman in a bikini → a dirty old manWoman in a bikini + the face→ a dirty old man
Blue: Kuleshov’s experiment as recounted by Pudovkin. Crimson: Hitchcock’s 1964 example.

Does it really work?

Science took its time to check. A 1992 study by Stephen Prince and Wayne Hensley, with 137 participants, failed to reproduce the effect, though it used a single trial that made it sensitive to noise. In 2006, Dean Mobbs and colleagues used brain imaging and found that neutral faces were read as sadder or happier depending on the scene before them. A 2016 study by Daniel Barratt and colleagues, with 36 participants across five emotions, confirmed it.

The effect is real, and subtler than the legend: context does not rewrite a face, it tilts how we read it. That tilt is what editing is made of.

Why it matters more with AI

A video model produces shots. Each one can be beautiful, and each one is made on its own: the model does not know what came before or what comes next. But meaning lives between shots.

That is why so many AI videos feel like a mood board in motion. A series of good images that add up to nothing, because nobody decided what the cut should make the viewer think. The same product shot can read as luxury after a quiet close-up of hands, as speed after a car at night, as comfort after rain on a window. The shot does not choose. The edit does.

Generation made shots cheap. It did not make meaning cheap. Even the famous one-take films are built on that decision.

A test for your next video

Take out one shot. If nothing changes in what the viewer understands or feels, it did not need to be there. Then swap two shots. If the meaning changes, someone is editing. If it does not, someone is decorating.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →