A brand that orders an AI video rarely needs a single shape. The website wants a horizontal 16:9 frame, the phone feeds want a vertical 9:16, and someone will suggest generating the wide version once and cropping the rest. Film history ran that experiment every time the screen changed shape, and each time the composition turned out to belong to one ratio.
Generated video sharpens the question, because a model composes each shot for the frame it is asked to fill.
A frame built for faces
Silent films were often shown at 1.33:1, the 4:3 shape, the U.S. National Archives notes. In 1932 the Society of Motion Picture Engineers adopted the Academy ratio, 1.37:1, for Hollywood sound films. For two decades, studio pictures were composed for a frame only slightly wider than it was tall, a shape that suits a face, a single figure or a doorway. Television later used the same 4:3 proportions. American Cinematographer still draws the line there: anything wider than the Academy aperture of 1.37:1 counts as widescreen.
Video models also compose for a shape, and the ratio picked at the start decides where the subject sits. In a February 2024 technical report, the developers of one model wrote that training on native aspect ratios improved framing, while a version trained on square crops “sometimes generates videos where the subject is only partially in view.”
CinemaScope, Hollywood’s wide answer to television
In the 1950s, American Cinematographer recounts, theatrical films faced off with broadcast television, and studios answered with widescreen, large formats, stereo sound, 3D and drive-ins. Fox’s CinemaScope used an anamorphic lens to squeeze a wide image onto ordinary 35 mm film, projected at 2.55:1, according to Britannica. The National Archives notes that the squeeze could fit an image up to 2.66:1 onto film that usually carried a 1.37:1 frame. Its first release, The Robe, premiered at the Roxy Theatre in New York on September 16, 1953. The magazine hailed the process as “the greatest development since the introduction of sound.” After the first week at the Roxy, Variety estimated that Fox could recoup up to a quarter of its investment from New York alone.
The wider frame changed how scenes were made. The AFI Catalog notes that CinemaScope required drastic changes in lighting and in the placement of actors within a scene, and director Henry Koster later recalled lenses that had to be focused separately and often forced retakes. The catalog also records that The Robe was shot in regular 35 mm as well, for a 16 mm version meant for churches and schools. In theaters it played in CinemaScope only, and by the end of 1953 it had grossed $16.5 million in 400 theaters, about a quarter of the U.S. theaters then equipped for the process. The film made to beat television became the top-rated show of its week when ABC broadcast it in March 1967.
A model asked for 16:9 often stages a scene across the width too: people at opposite thirds, a product on one side. Matting that frame to 2.39:1 keeps three quarters of its height. A vertical crop keeps far less.
Shoot and protect, the compromise behind 16:9
Wide films eventually reached 4:3 television. Mark Schubin’s 1996 tutorial in the SMPTE Journal describes the usual method: match the heights of the two images and remove the excess width. How to Marry a Millionaire, made for 2.55:1 projection, was cut down to television’s 4:3 in 1961. Scope itself narrowed as prints carried an optical soundtrack, and a 1971 revision of the projection standard trimmed the aperture to hide splices, leaving about 2.39:1. Flat widescreen took another route: Schubin notes that 1.85:1 is the projection ratio most commonly recommended in the U.S. for non-anamorphic 35 mm films.
High-definition television was designed to end the war of shapes. Schubin recounts that Kerns Powers, then at RCA Laboratories, presented a proposal to an SMPTE working group on May 4, 1984. He took the narrowest common shape, 4:3, and the widest, 2.35:1 scope, and derived a frame between them: just over 1.77:1, rounded to 16:9. His drawing was a shoot-and-protect shape, a method Schubin defines as framing for one ratio while protecting extra area so the image still works in another, with no lights, microphones or set edges showing. In practice, reference lines in the viewfinder mark the action area for one ratio inside the frame for the other.
A century of screen shapes
Each ratio drawn at the same height, with the moment it took hold
| Ratio | Name | Context |
|---|---|---|
| 1.33:1 | Silent 4:3 frame | Silent era |
| 1.37:1 | Academy ratio | Adopted in 1932 for sound films |
| 2.55:1 | CinemaScope | 1953, The Robe |
| 2.39:1 | Scope | Projection standard revised in 1971 |
| 1.78:1 | 16:9 | Proposed in 1984 for high-definition television |
| 0.56:1 | 9:16 vertical | Phones; IGTV launched vertical-only in 2018 |
AI projects need that method when one master must serve several screens. A 9:16 window cut from a 1,920 by 1,080 frame is about 608 pixels wide, under a third of the width, and must be enlarged almost 1.8 times to fill a phone. A subject on a third gets sliced. Protecting the frame means a first frame with the face and product in the central column, and every move checked against it.
The shape is part of the story
Filmmakers still pick ratios for meaning. Quebec director Xavier Dolan presents most of Mommy in a 1:1 square, which critic Glenn Kenny reads as trapping the viewer in the hemmed-in options of its characters. In a montage set to Wonderwall, as Steve, Diane and Kaya start to blossom, Steve “almost literally ‘opens’ the frame” and the film breathes in widescreen. Dolan widens it once more for a dream sequence in which an idealized Steve builds a family, then returns to the square, a move Kenny finds cruel to Diane and to the audience.
Phones brought the opposite of CinemaScope. Instagram launched IGTV in June 2018 as “built for how you actually use your phone,” with full-screen vertical video. A 9:16 frame is 16:9 turned on its side: it holds one standing person, a face above a product, elements read from top to bottom. It has little room for two people side by side, which is what a wide frame does best.
Generating natively in 9:16 lets the model compose for that shape, at a cost in consistency: each generation starts fresh, so the vertical version can come back with a different face, jacket or light. The fix is the one used between shots: the same approved references, and a keyframe in each ratio compared side by side before animation.
The third path, outpainting, invents image beyond the edges: turning 16:9 into 9:16 at the same width means inventing about two thirds of the frame. A March 2024 paper by Fu-Yun Wang and colleagues on fitting videos to phone screens names the hard part, keeping new content consistent within and across frames. Invented sky can pass; a logo cannot.
What to ask before the first generation
Ask which ratios the project delivers and whether each version is generated natively, cropped from a master or outpainted. Approve a keyframe in every ratio, side by side, and check that faces, wardrobe, product and light match. In every crop, follow your product, logo and on-screen text from first frame to last, and keep your brand out of any invented edge. Fox’s crews restaged scenes for a new frame in 1953; a vertical ad deserves the same care.
Sources
- The Measure of a Screen: Motion Picture Aspect Ratios in the Archives (Heidi Holmstrom, The Unwritten Record, U.S. National Archives, January 18, 2018)
- What Is Widescreen? (David E. Williams, American Society of Cinematographers, February 27, 2025)
- The Robe, film by Koster (Encyclopaedia Britannica)
- The Robe (1953) (AFI Catalog of Feature Films, American Film Institute)
- Searching for the Perfect Aspect Ratio (Mark Schubin, SMPTE Journal tutorial, August 1996)
- Anamorphic format (Wikipedia, citing the SMPTE projection standard of June 1971)
- Mommy (Glenn Kenny, RogerEbert.com, January 23, 2015)
- Welcome to IGTV (Instagram, June 20, 2018)
- Video generation models as world simulators (OpenAI technical report, February 15, 2024)
- Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation (Fu-Yun Wang et al., arXiv, March 2024)