A fifteen-second generated ad can look effortless. Behind it sits a list of problems every AI filmmaker meets on every project, whatever the model: faces that change, labels that melt, sound that has to be rebuilt, rights that have to be cleared.
None of them is a reason to avoid generated video. Each has a cause, and each has a response that working professionals now treat as routine. Here are the ten that matter most to a brand, grouped by where they show up.
In the frame: what the model draws
1. The same face, shot after shot
Each generation starts fresh. The model keeps no memory of the previous shot, so a face, a jacket or a coffee cup is redrawn from whatever description and images it receives. When the Toronto studio Shy Kids made the short film Air Head in 2024 with an early-access video model, its balloon-headed hero would not stay yellow from shot to shot. The team isolated and recoloured the balloon in post, Mike Seymour reported for fxguide.
Professionals lock identity before production: a character sheet, the same reference images in every generation, and a still from the last approved shot as the first frame of the next. The method is laid out in Character Consistency in AI Video, and props drift for the same reason, as the script supervisor’s job shows.
2. Products, text and logos
A model draws letters as shapes it has seen, without spelling them. A label, a price or a logo is one more texture to approximate, so packaging often comes out almost right, the worst result for a brand. T2VTextBench, a human-rated benchmark from 2025, tested ten video models and found that “most struggle to generate legible, consistent text.” VidScribe, a benchmark posted in September 2026, found text holds up best when a generation starts from an image, and worst when existing video is edited.
The working answer is to keep the real product in charge: start from an approved packshot as the first frame, keep the label large and steady, and add logos, supers and end cards in the edit instead of generating them.
3. Weight, physics and hands
Video models learn what motion looks like, not why it happens. In a 2024 study, researchers led by Bingyi Kang trained video models on simple simulated physics. In familiar situations the models predicted motion well. Outside them they failed, “mimicking the closest training example” instead of applying a rule. Hands add their own trouble, with many joints often hidden behind the object they hold. A hand gripping a bottle is where fingers merge.
The response is to choose actions the model has seen countless times, keep each movement simple, review shots frame by frame, and film real hands when close product handling matters. The visible symptoms are listed in Why Your AI Video Looks Like AI Video.
In motion: performance, sound and length
4. Performance and lip sync
A generated face can move correctly and still read empty, and a mouth has to land every consonant on time. The tolerance is small. Recommendation BT.1359 of the International Telecommunication Union, from 1998, puts the detection threshold at about 45 milliseconds with sound ahead of picture and 125 milliseconds with sound behind it. At 24 frames per second, one frame lasts about 42 milliseconds.
So the voice comes first, recorded or designed, and the picture is generated to it. Sync is checked frame by frame, and scenes are written to need less of it: voice-over, a speaker in profile or off screen. Why near-human faces unsettle viewers is covered in Why Almost-Human Faces Make Audiences Flinch.
5. Sound
Many generated clips arrive silent, and models that do generate audio produce it shot by shot, so the room tone changes at every cut. Viewers hear those seams. The fix is to treat sound as a production of its own: one ambience bed across the sequence, foley, a real mix and licensed music, keeping generated audio only where it survives that mix. The case is made in AI Video Has a Sound Problem.
6. Short clips, long ads
A single generation runs from a few seconds to about half a minute at most, and the longer it runs, the more the image can drift. A 30-second spot is still a sequence of shots that must cut together. Professionals plan the edit before generating anything: a shot list with durations, extra frames at each end for trimming, cuts placed on motion, and seams hidden the way long-take films have always hidden them, as The One-Take Shot Is Cinema’s Most Beautiful Lie explains.
In the budget: selection and revisions
7. The selection ratio
A director cannot ask a model for one more take with a little less smile. Each generation is a new draw, and most are discarded. For Air Head, Shy Kids’ Patrick Cederberg estimated about 300 units of generated material for every unit in the final film, according to fxguide. Every rejected clip costs compute, and someone’s time to review it.
Good teams track that ratio by type of shot and spend it carefully: cheap drafts to test motion, still frames approved before anything is animated, and final-quality renders only for shots already locked.
8. Revisions that restart the shot
On a shoot, “a bit warmer” is a colour-grading note. In generated video, changing the prompt reruns the whole shot, and the new version can lose details the client had already approved. The tool can also change mid-project. Google’s cloud documentation says stable model versions are typically released with a retirement date, after which “API requests to a retired model will fail.” A major video product has already gone dark.
Professionals approve in stages (script, first frames, motion, finishing), freeze what is approved, archive every prompt, reference and model version, and define in writing what counts as a revision round.
In public: rights and trust
9. Ownership and consent
Copyright follows human authorship. The U.S. Copyright Office concluded in January 2025 that prompts alone do not make someone an author, while human selection, arrangement and modification of generated material can be protected. On March 2, 2026, the U.S. Supreme Court declined to hear Thaler v. Perlmutter, leaving the human-authorship requirement in place. In Canada, a February 2025 federal consultation report noted that case law points to authorship by “a human who exercises skill and judgment,” and announced no change to the law.
Faces and voices need their own consent. The SAG-AFTRA Commercials Contracts, in effect since April 1, 2025, require a performer’s clear written consent to a digital replica, separate from the employment contract, with a reasonably specific description of its use. Professionals document the human creative work, use references they own, avoid prompting real people or protected characters, and get written consent for any real face or voice. More in What Makes an AI Spokesperson Believable.
10. Disclosure and trust
Some markets now require a label, as in the European Union, and the evidence on what a label costs is mixed. A study in the Journal of Retailing and Consumer Services found that labelling an ad as AI-generated lowered attitudes toward it through a sense of eeriness, an effect that faded for a stronger brand. A MediaScience test of four label formats with 900 U.S. participants, reported by Marketing Dive in June 2026, found labels did not lower brand recall. It also found that even a label shown for the whole ad left only 49% of viewers aware AI was used. An icon reached 38%.
Professionals settle disclosure in the brief, prefer plain text to an icon, and never let a generated person pose as a real customer.
What a brand should check
Most of these problems are prevented before generation or caught after it. Few are solved by the model alone.
Where each challenge gets handled
The ten challenges mapped to four stages of an AI video project
| Stage | Challenges handled |
|---|---|
| Prepare | 1 Same face, 2 Product and text, 4 Lip sync, 6 Clip length, 8 Revisions, 9 Rights and consent |
| Generate | 3 Physics and hands, 7 Selection ratio |
| Edit and sound | 2 Product and text, 5 Sound, 6 Clip length |
| Review and release | 1 Same face, 3 Physics and hands, 4 Lip sync, 8 Revisions, 10 Disclosure |
Before approving an AI video project, ask the team which of the ten it expects to face, then check:
- the character and product references, and the first frames, before any motion is generated;
- where the real product, logo and on-screen text come from, and whether they are added in the edit;
- how revision rounds are counted, and whether prompts, references and model versions are archived;
- the written consent for any real face or voice, and the disclosure plan before launch.
Sources
- Making the short film Air Head with an AI video model (Mike Seymour, fxguide, April 14, 2024)
- T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models (Guo et al., arXiv, May 2025)
- Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation (Zhang et al., arXiv, September 2026)
- How Far is Video Generation from World Model: A Physical Law Perspective (Kang et al., arXiv, November 2024)
- Recommendation ITU-R BT.1359: Relative timing of sound and vision for broadcasting (International Telecommunication Union, 1998)
- Model versions and lifecycle (Google Cloud documentation)
- Copyright Office Releases Part 2 of Artificial Intelligence Report (U.S. Copyright Office, January 29, 2025)
- The Final Word? Supreme Court Refuses to Hear Case on AI Authorship and Inventorship (Jacob W. S. Schneider, Holland & Knight, March 3, 2026)
- Consultation on Copyright in the Age of Generative Artificial Intelligence: What We Heard Report (Innovation, Science and Economic Development Canada, February 11, 2025)
- New SAG-AFTRA Commercials Contract Is Now in Effect (Davis+Gilbert LLP, 2025)
- AI disclosure and advertising effectiveness: The uncanny valley of mind and the buffering role of brand strength (Mo, Zhu, Wu and Deng, Journal of Retailing and Consumer Services, vol. 94)
- AI disclosure labels don’t hurt ad performance: Here’s what the numbers say (Marketing Dive, June 12, 2026)