In late March 2023, a Reddit user posted a montage of Will Smith eating spaghetti. Nothing in it was filmed. A free research model had generated fragments of one or two seconds, which the poster edited together, Vice reported, with a stock-photo watermark still visible across the frames. The face melted, the noodles fused with the hands, and the whole thing was silent.

Three years on, brands are being offered AI video for real campaigns. A marketing lead deciding whether to buy needs a clear record of what the models gained since that clip, and of the work that still falls to people. Here it is, dated and sourced.

REC 2023 TO 2026 2023 2 S NO SOUND 2026 30 S SOUND LONGER, SHARPER, WITH SOUND. STILL ONE SHOT AT A TIME.
Three years of gains, all measured inside a single shot.

The clip that became a yardstick

The spaghetti montage was typical of its moment. Gizmodo’s survey of text-to-video in March 2023 described the open research model behind it as making two-second clips, and a commercial tool, then only in demo, as making three. When that tool opened to the public in June 2023, Tom’s Hardware found it produced four-second clips with “no sound at all” and very limited movement. Higher-resolution versions were reserved for paying accounts.

The clip outlived its technology. Each time a new video model launched, people ran the same prompt on it, and by the end of 2024 TechCrunch called the exercise “something of a meme as well as a benchmark.” Smith himself parodied it on Instagram in February 2024. The test works because it is hard: one face that must stay the same person, two hands, a fork, and food that changes shape as it is eaten.

What the models learned in three years

Length moved first. In February 2024, OpenAI published research on a model it said could generate a minute of high-fidelity video, at up to 1920x1080 pixels. In July 2026, ByteDance’s Seed research team announced a model that generates single shots of up to 30 seconds, double its previous 15-second limit, and that can extend them over several rounds.

Sound arrived in May 2025, when Google introduced a model that generates audio with the picture: traffic noise, birdsong, even dialogue between characters, with what the company described as accurate lip sync. Leading models released since then generate sound too: the July 2026 model produces audio and video jointly, and Google’s current developer documentation lists audio as always on.

Control matters most for brands, and it came last. A 2023 prompt was text and hope. Current models accept reference inputs: images of a product, a face or a set that the generation has to respect. Google’s documentation allows up to three reference images per clip. The July 2026 model accepts up to 30 images, 10 video clips and 10 audio clips in a single request. Resolution climbed along the way. In 2023, the commercial tool charged extra for an upscaled file; the same Google documentation now lists 4K output.

Three years of AI video, dated

Capability milestones from published reports and vendor documentation

MAR 2023MAR 2023: Viral clip, free research model, 1 to 2 s fragments, edited together No sound, stock watermarkViral clip, free research model1 to 2 s fragments, edited togetherNo sound, stock watermarkJUN 2023JUN 2023: Commercial tool opens to all, 4 s clips, no sound Upscaling for paid accountsCommercial tool opens to all4 s clips, no soundUpscaling for paid accountsFEB 2024FEB 2024: Lab research results, Up to a minute of video Up to 1920x1080Lab research resultsUp to a minute of videoUp to 1920x1080MAY 2025MAY 2025: Sound generated with picture, Effects, ambience, dialogue Lip sync claimed accurateSound generated with pictureEffects, ambience, dialogueLip sync claimed accurateJUL 2026JUL 2026: Single shots of up to 30 s, Up from 15 s; references: 30 images, 10 clips, 10 audioSingle shots of up to 30 sUp from 15 s; references: 30images, 10 clips, 10 audioOCT 2026OCT 2026: Leading model, current docs, 8 s shots up to 4K, sound on Extensions to 148 s, at 720pLeading model, current docs8 s shots up to 4K, sound onExtensions to 148 s, at 720p
AI video capability milestones, 2023 to 2026
DateMilestoneClip lengthResolutionSound
March 2023Viral Will Smith clip from a free research model1 to 2 seconds per fragmentNot statedNone
June 2023Commercial text-to-video tool opens to all4 secondsUpscaled version for paid accountsNone
February 2024Lab research resultsUp to one minuteUp to 1920x1080Not described
May 2025Model generates sound with pictureNot statedNot statedEffects, ambience, dialogue, lip sync
July 2026Single-shot model with many referencesUp to 30 seconds (from 15)Not statedGenerated jointly with video
October 2026Leading model’s current documentation4, 6 or 8 seconds; extensions up to 148 secondsUp to 4K (8-second shots); 720p for extensionsAlways on
Sources: Vice and Gizmodo (March 2023), Tom’s Hardware (June 2023), OpenAI (February 2024), Google (May 2025), ByteDance Seed (July 2026), Google AI for Developers documentation (consulted October 1, 2026). Vendor figures are the vendors’ own claims.

The fine print on longer and sharper

Spec sheets make the problem sound solved. The documentation is more careful. On Google’s current model, a single generation still runs 4, 6 or 8 seconds, and 4K is only offered at 8 seconds. Longer pieces come from extension: the model adds seven seconds at a time, up to 20 times, for a maximum of 148 seconds, and only at 720p.

So the length a model can reach and the quality it holds at that length are two different numbers. A 30-second single shot is real progress for a scene that needs one continuous move. A brand film is still built from many shots, and each one is a separate generation that has to match the others.

What three years did not fix

Story is still written by people. Every milestone above concerns a single shot: longer, sharper, with sound. None decides what the ad says, which shot comes first or where the cut lands. A model will render a person eating for as long as you ask. Someone still has to decide why the audience should keep watching.

Continuity across shots is still a direction job. Among its own model’s limits, OpenAI noted in 2024 that eating food did not always change the food correctly. Inside a single shot, models have improved a great deal since then. Between separately generated shots, a face, a product label or the light can still shift. References reduce that drift; someone still has to check every cut.

Sound judgement did not arrive with sound. When users ran the spaghetti test on Google’s new model in May 2025, PetaPixel described a convincing Smith making “crunching and squelching noises” while chewing pasta. The audio was synchronized and wrong for the scene. Deciding what the viewer hears, what is left out and how the product sits against the voice remains a mixing decision, and part of the generated audio usually gets replaced in post.

Rights moved, and not toward shortcuts

The 2023 clip used a famous face that nobody had licensed. The models have changed and that question remains. In February 2026, Axios reported that Disney had sent ByteDance a cease-and-desist letter over its new video model, accusing it of reproducing and creating derivative works from Disney characters without permission.

Ownership of the output is unsettled as well. In January 2025, the U.S. Copyright Office concluded that “the mere provision of prompts” does not make someone the author of what a model produces. Protection can still apply where a person determines enough of the expressive elements, for instance by modifying the output or arranging it into a larger work. For a brand, the human contribution to an AI video becomes something worth documenting.

What a brand can expect now, and must still check

A brand can now expect clean single shots of several seconds, synchronized sound, higher resolution and a product or face held close to its reference within each shot. It should still check:

  • The resolution each final shot was generated at, and whether long shots were extended at lower resolution, then upscaled.
  • How the product, faces and set were kept consistent from one shot to the next.
  • Who designed the sound, and which parts of the generated audio were kept.
  • Which references were used, who owns them, and whether any real person, character or brand appears without permission.
  • What human creative work is documented, in case ownership of the film is ever questioned.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →