A brand that wants an AI spokesperson can generate a face in minutes that passes for a photograph. The trouble starts when it talks. A blink lands late, the teeth blur together, a smile arrives a beat too slowly, and viewers stop trusting what they see without knowing why.

Film has a name for that feeling, and paid dearly to learn it. Two decades ago, studios spent fortunes on computer-generated humans that audiences found creepy. Their lessons about eyes, motion and style apply almost point for point to the avatars brands commission today.

REC THE UNCANNY VALLEY STILL MOVING HUMAN LIKENESS AFFINITY STYLISED ALMOST HUMAN HUMAN
Mori’s curve: movement deepens both the peak and the valley.

A roboticist’s graph, 1970

In 1970, the Japanese roboticist Masahiro Mori published a short essay in a journal called Energy. IEEE Spectrum ran the first English translation authorized and reviewed by Mori in June 2012. His idea fits in one graph. As a robot looks more human, our affinity for it rises, until the likeness gets close but not quite right. There the curve drops into what he called the uncanny valley, then climbs again toward a healthy person.

He placed the bunraku puppet of Japanese theatre high on the curve: seen from the seats, its stylised face reads as human and earns real affinity. A realistic prosthetic hand, by contrast, looks convincing until you shake it. Movement amplifies both the peaks and the valleys. He described a robot face built to smile; when its smile was slowed to half speed, the expression “turns creepy.” His advice to designers was to aim for the first peak, a moderate human likeness, and skip the risky climb toward full realism.

That slowed smile is a familiar failure in generated video. A model can render a still face that would fool anyone, then animate it with timing no face has: a smile that takes two seconds to form, a blink that half closes, features that slide as the head turns. Frame interpolation makes it worse, since every in-between frame is invented. Judge an avatar in motion, at full speed, never from the hero still.

The first photoreal cast

Square Pictures tried to climb the whole curve in one film. Final Fantasy: The Spirits Within, released in July 2001, built its human cast entirely in computer animation, led by a scientist named Aki Ross. Roger Ebert, who gave the film two stars, admired the craft and still drew the line where Mori would have: “Not for an instant do we believe” she was a real human, he wrote, while conceding she was lifelike. Her movements, taken from real actors, felt about right to him, and her hair moved convincingly in the wind.

Audiences stayed away. Variety reported a budget of $145 million against a worldwide gross of $72 million, for an estimated loss of about $73 million. Square shut its five-year-old Honolulu studio on March 31, 2002, affecting 125 employees.

Generated avatars fail more cheaply, but the same way. A demo shows a striking face for three seconds, and the brand approves a 60-second script. AI video models generate a few seconds at a time, so a long talking head is stitched from separate generations, and the face, the light or the voice can shift at every seam.

The Polar Express and the missing eyes

Three years later, Robert Zemeckis’s The Polar Express took another route to realistic humans: performance capture, with performers in bodysuits covered in infrared sensors. Tom Hanks played five roles, including the conductor. Reviewing it for CNN in November 2004, Paul Clinton found the human characters creepy. He traced it to the eyes and the inside of the mouth, which the sensors could not reach and which had to be computer-generated. The result, he wrote: “these characters look soul dead.” When they stood still, they looked lifeless.

Research later backed his instinct. In a 2011 study in Computers in Human Behavior, Angela Tinwell and colleagues found that even fully and expertly animated characters were rated more uncanny than humans, and that a lack of movement in the upper face during speech made it worse, especially for fear, sadness, disgust and surprise.

Generated talking heads share the weak spot. Audio-driven lip sync puts its effort where the sound is, the mouth, while brows, eyelids and cheeks can stay oddly still. The usual tells: gaze that never lands on the lens, blinks on a metronome, teeth that fuse or change shape between words, skin with no pores. Together they reproduce what Clinton saw in 2004.

Where a generated face slips into the valley

Five zones to check on any AI avatar, in motion

12345Eyes: Gaze that never lands, blinks on a metronome1EyesGaze that never lands,blinks on a metronomeUpper face: Brows frozen while the mouth talks2Upper faceBrows frozen whilethe mouth talksMouth: Lips that miss consonants, teeth that blur or shift3MouthLips that miss consonants,teeth that blur or shiftSkin: Too smooth, no pores, light that doesn’t match4SkinToo smooth, no pores,light that doesn’t matchTiming: Smiles that unfold slowly, features that drift5TimingSmiles that unfold slowly,features that drift
Five zones where a generated face slips into the uncanny valley
ZoneWhat to look for
EyesGaze that never lands on the lens, blinks at a mechanical rhythm
Upper faceBrows and eyelids frozen while the mouth talks
MouthLip shapes that miss consonants, teeth that blur or change shape
SkinToo smooth, no pores, light that does not match the scene
TimingSmiles that unfold too slowly, features that drift during a head turn
Based on Mori (1970, English translation 2012), Paul Clinton’s review of The Polar Express (CNN, 2004) and Tinwell et al. (Computers in Human Behavior, 2011), applied to generated video.

Two ways out: style or direction

Pixar learned the lesson early. Its 1988 short Tin Toy featured a realistic baby that some viewers called “demonic,” as Animation Obsessive recounts, and Ed Catmull later wrote that the studio did not want its first feature to have humans as the major characters.

For the 1997 short Geri’s Game, director Jan Pinkava turned to European puppet theatre and designed his old chess player as a digital puppet. When humans took the lead in features, Pixar stylised them too. Reviewing the trailer for Brad Bird’s The Incredibles in October 2004, Cartoon Brew’s Amid Amidi praised human characters “that didn’t try to emulate life.” That film and The Polar Express came out the same year, from opposite ends of Mori’s curve.

Generated faces complicate the picture. In a 2022 PNAS study by Sophie Nightingale and Hany Farid, 315 people trying to tell real face photos from AI-synthesized ones scored 48.2%, close to chance. The authors concluded that synthesis engines have passed through the uncanny valley. But those were still images. An avatar has to move and speak, which puts it back on Mori’s moving curve.

Audiences online still react as Mori predicted. Over 11 months, Jbid Arsenyan and Agata Mirowska compared reactions to three accounts on a popular social platform: a human influencer, a human-like virtual one and an anime-style virtual one. The human-like virtual account drew significantly fewer positive reactions.

That leaves a brand two honest routes. A clearly stylised avatar sits on Mori’s first peak, like the bunraku puppet. A realistic one has to be directed like a performance: a real actor’s take used as motion reference or run through video-to-video so the timing is human, a locked reference sheet for the face, short shots, lip sync checked word by word. A realistic face left to a model’s default motion is the valley.

Test it on people who didn’t make it

The team that built an avatar has watched it hundreds of times; fresh viewers are harsher. In NielsenIQ’s December 2024 research, more than 2,000 people watched AI-generated ads, about 150 of them under EEG. They readily spotted the AI ads, called them “annoying,” “boring” and “confusing,” and even high-quality ones produced weaker memory activation.

A small test catches most of the valley before launch. Show the avatar to five or ten outsiders, on a phone, at full speed, without saying it was generated. Ask what felt off. Play it again with the sound off, so a good voice cannot cover weak lip sync. Torn between realistic and stylised? Test both.

Then decide how to disclose. In a 2025 study, Jamie Grigsby, Meg Michelsen and César Zamudio found that AI disclosures in service ads lowered trust, which recovered when AI generated only tangible elements, such as an office, and the service provider stayed real. So choose what you generate, disclose it plainly, and never let viewers discover a realistic synthetic person on their own.

What to ask of any AI avatar project

Ask to see the avatar moving, for the full length of the script, before approving its look. Ask whether it should read as stylised or real, and if real, whose performance drives it. Check eyes, brows, teeth and skin in close-up, once with the sound off. Ask which outsiders have seen it. Agree on how viewers will learn it was generated.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →