Ask a video model for your bottle on a sunlit counter and you will usually get something close: the right colour, roughly the right shape, and a label whose letters swim as the bottle turns. For a brand, close is the problem. A viewer forgives a generated kitchen. A shopper who finds a different label on the shelf does not.
Commercials faced a version of this long before AI. Tabletop crews, the specialists who shoot food, drinks and packaged goods, spent decades learning to make a product look perfect while keeping it the product. Their rules are a good guide to what generated video can take over and what it should leave alone.
Fake the milk, never the cereal
Delores Custer, a food stylist with more than 30 years in the trade, described her toolkit to NPR in 2010. Chicken is undercooked, then painted with angostura bitters, yellow food colouring and a dash of detergent for a roasted gloss. A perfect splash of milk might be sculpted acrylic or even Elmer’s glue. Cutlery that slides is held in place with mortician’s wax.
Drinks have their own kit. Atlas Obscura reported in 2015 that stylists use plastic ice when real ice melts too fast under the lights, and spritz glasses with corn syrup and water so the drops bead like condensation.
The item being sold is another matter. “It’s really important to tell everyone that it’s real food,” stylist Janine Kalesis told CNBC in 2014. In the same report, stylist Nir Adar said that working for Burger King, he had to sign a legal document stating he hadn’t altered anything.
The line has a history. In 1970, Britannica recounts, the Federal Trade Commission took action against Campbell Soup over a vegetable soup commercial whose photographers had put clear marbles in the bowl to push the vegetables to the surface.
A video model knows no such line. It repaints the whole frame, product included, so the cereal is as negotiable as the milk. The rules did not loosen for it: the UK’s Committee of Advertising Practice wrote in 2023 that its code applies “regardless of how the ad was created.” A generated scene can invent the counter and the morning light. What a shopper will compare with the shelf has to come from the real thing.
The hero is chosen, then documented
On a shoot, the hero is whatever is starring in the shot, in Atlas Obscura’s definition. Stylists earn it by selection. The Big Mac in a McDonald’s ad really is a Big Mac, NPR noted, assembled from the most photogenic of thousands of buns, patties and lettuce leaves.
In June 2012, McDonald’s Canada took on the question of why its food looks different in ads. In a video presented by marketing director Hope Bagozzi, a team rebuilt a Quarter Pounder with Cheese from the ingredients its restaurants use. As NPR reported, the condiments were moved forward in the frame, tiny holes in the bun were fixed and the cheese was retouched into an even melt. One burger took hours.
In AI work, the hero becomes a reference set: photographs of the exact SKU in its current packaging, with the right flavour, size and language version. Without it, the model returns its average bottle. In March 2025, Unilever said it builds physically accurate 3D replicas of its products holding “all its variants, labels, packaging and language formats” in one file, and credited them, by its own figures, with imagery made twice as fast at half the cost.
Macro, high speed and a robot that repeats itself
Tabletop cinematography stretches time and space. The Phantom Flex4K high-speed camera records 1,000 frames per second at 4K, according to its maker, so one second of a pour plays for about 42 seconds at 24 frames per second. Under a macro lens, a sesame seed fills the frame, and a centimetre of camera travel becomes a large move.
Hence the robot arm. Mark Roberts Motion Control’s Bolt, sold for “live-action or tabletop shoots,” reaches two metres, carries a 20 kg camera, moves it at up to 7 metres per second and repeats each programmed move precisely. Its maker says it goes from a standstill to full speed and back to a stop in fractions of a second, so a fast move can start and stop on cue. A crew can run the identical move several times and layer the passes, including a clean pass with nothing in frame for compositing. Rental house General Lift describes using such rigs to move the food or product and the camera “in concert.”
Generated video has no repeat button. The same prompt run twice gives two different moves, and a slow-motion pour comes out as a plausible liquid, not yours. A first or last frame, a motion reference from a real camera move, or a video-to-video pass over filmed footage narrows the drift without removing it. Anything that must match from shot to shot, starting with the product, is safer captured once and reused.
Lighting a label is its own job
Glass and print want opposite light. In a 2024 guide to wine bottle photography, Laurie Millotte of Outshinery recommends a soft source at 30 to 60 degrees so the highlight runs down the side of the glass, and a separate small, narrow source aimed at the label to keep the type crisp. “Fix the lighting for one and you usually break another,” she writes. Black cards and flags block unwanted reflections.
The label is where generated video is weakest. A benchmark posted on September 29, 2026, by Ziying Zhang and colleagues tested 11 video systems on 803 human-verified samples of on-screen text. Getting the words right and drawing each letter correctly proved to be separate skills. Text held best when generation started from a supplied image, and editing text inside existing video was the main bottleneck. So start from a real photograph of the product, keep moves around the label small, and treat any text the model writes on its own as wrong until someone reads it.
Film the product, generate the world
The safest hybrid splits the frame. The product is photographed or filmed for real, lit by someone who knows how to light a label, and the surroundings are generated around it. A compositor then matches what ties the two together: light direction and colour, lens perspective, the contact shadow under the product and the reflections in its glass. When those agree, the eye accepts the scene. When they don’t, the product looks pasted on.
Who owns each layer of a hybrid product shot
The product is captured; the world around it can be generated
| Layer | Source |
|---|---|
| Environment: counter, room, weather | Generated |
| Light direction and colour | Matched to the product plate |
| Contact shadow and reflections | Matched in compositing |
| Product: label, logo, shape, fill level | Filmed or photographed from the exact SKU |
Some shots should stay fully captured. Anything that demonstrates what the product does, such as the head on a beer, the stretch of melted cheese or the texture of a cream, is a claim about the product and belongs to the product. So is a close-up meant to be read, and any shot that shows real size against a hand.
What to ask before you approve
Ask where every product pixel comes from: a photograph, a filmed plate, a 3D replica or a generation. Ask for the reference set of the exact SKU and check that it shows the current packaging and the right market version, French label included. Review the final cut frame by frame at full resolution, reading every letter of the logo, the label and the claims, and compare shape, fill level and colour with a unit taken off the shelf. Then ask which shots show the product performing. Those should be filmed with the real thing, the way the Big Mac in the ad is a Big Mac.
Sources
- Food Stylists Put The Sizzle On Your Burger (NPR, May 22, 2010)
- Why restaurant meals don’t look like the ads (Katie Little, CNBC, May 19, 2014)
- Advertising fraud: Television and consumerism (Encyclopaedia Britannica)
- Generative AI & Advertising: Decoding AI Regulation (CAP News, Advertising Standards Authority, August 31, 2023)
- Pizza Doesn’t Really Stretch That Way, and Other Devious Food Stylist Tricks (Cara Giaimo, Atlas Obscura, August 21, 2015)
- Watch McDonald’s Make A Supermodel Out Of A Quarter Pounder (The Salt, NPR, June 22, 2012)
- Unilever reinvents product shoots with digital twins and AI (Unilever press release, March 18, 2025)
- Phantom Flex4K specifications (Vision Research)
- Bolt high-speed cinema robot specifications (Mark Roberts Motion Control)
- About motion control shots (General Lift)
- Wine bottle photography: shoot like a pro (Laurie Millotte, Outshinery, December 2024)
- Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation (Ziying Zhang et al., arXiv, September 29, 2026)