On October 6, Google released Nano Banana 2.1, the new version of the image model behind the Gemini app. Google’s own line is that it “outperforms our previous models across the board, with notable leaps in visual design, mask-based editing, and subject consistency.” For developers, the headline is price: a standard 1K image now costs half what it did on Nano Banana 2.
For a studio that makes AI video, the image model is the first link in the chain. Character sheets, location plates, storyboard frames and product stills are generated as images first, then animated. A cheaper, more consistent image model changes the cost and the reliability of everything downstream. Here is what shipped, what the independent votes say, and where it still falls short.
What Google shipped
In the Gemini API the model is called gemini-nano-banana-2.1. Google’s model card lists it in the Gemini app, AI Mode in Search, Google Ads, Google AI Studio, the Gemini API, Flow and Stitch. The developer documentation lists the main capabilities: output at 1K, 2K or 4K, up to 14 reference images in one request, grounding in Google web and image search, video as an input to produce a still, and three thinking levels (minimal, medium and high, where Nano Banana 2 had only two). Every image it produces carries a SynthID watermark.
The editing changes are the ones that matter most for production. Mask-based editing lets you mark one area of an image and change only that area. Subject consistency means the same character or product should look the same across a series of edits. Both address the most common reason an AI image gets thrown away: one wrong detail in an otherwise good frame.
The price, and the catch on the input side
Google’s pricing page lists $0.0336 per 1K image, against $0.067 for Nano Banana 2. A 2K image drops from $0.101 to $0.0504, and a 4K image from $0.151 to $0.113. Batch processing, for jobs that can wait, costs half again: $0.0168 per 1K image.
Half the price at 1K and 2K, a quarter off at 4K
Gemini API standard price per image, in US dollars
| Resolution | Nano Banana 2 | Nano Banana 2.1 |
|---|---|---|
| 1K image | $0.067 | $0.0336 |
| 2K image | $0.101 | $0.0504 |
| 4K image | $0.151 | $0.113 |
The input side moves the other way. Text, image and video input now costs $1.50 per million tokens, against $0.50 on Nano Banana 2. Most single prompts are small, so the output saving wins. A workflow that sends many large reference images with every request should check its own bills: the more references per call, the more the input price counts.
A concrete case: a storyboard of 40 frames at 2K, each generated three times before one is kept, is 120 images. That is about $12 on Nano Banana 2 and $6 on Nano Banana 2.1, before input costs. Small numbers per project, but they decide how many alternatives a team can afford to look at.
Google’s scores against the blind votes
Google’s model card reports an overall text-to-image preference score of 1050 (±14), an Elo-style rating from human comparisons. Decrypt reports Google’s comparison points of 990 for Nano Banana 2 and 935 for Nano Banana Pro. The largest gain in the card is multi-character consistency in editing, at 1106. For infographic factuality, Google reports 0.521, against 0.179 for the previous model according to Decrypt.
Those are Google’s tests. The first outside check comes from Arena, where users vote between two anonymous images. On October 6, its text-to-image leaderboard placed Nano Banana 2.1 fifth, at 1328 (±9) from 5,312 votes, a preliminary rating. The four models ahead of it were two GPT Image 2.5 variants, GPT Image 2 and Microsoft’s MAI-Image-2.6. On the image editing board, it sat sixth at 1428, level with MAI-Image-2.6, with the top OpenAI entry at 1524.
The two results do not contradict each other. Nano Banana 2.1 is clearly better than Google’s previous models, and clearly not the top image model on the market. For a buyer, the useful reading is that it is now a strong mid-price option, not the default answer for every job.
What Google says it still gets wrong
The model card is candid. Small text is often blurry. Character consistency between an input image and the output is imperfect. Edits guided by rough sketches follow instructions only partly. World knowledge and 3D reasoning are limited, and the knowledge cutoff is March 2026. In practice, that means a product label, a phone screen or a sign in the background still needs checking at full size, and a character sheet still needs a human eye before it feeds a video model.
What it changes for AI video production
Reference sheets get cheaper to iterate. When a character, a prop or a location has to stay identical across a whole film, most of the work happens before animation: generating clean sheets, comparing them, fixing details. Halving the price of that stage pays for more variations and a better final choice.
Fixes get smaller. Mask-based editing means a wrong jacket colour or a misplaced logo can be corrected in place, instead of regenerating a frame whose composition was already approved. For a client review, that is the difference between one revision and a new round.
Scenes with several characters get more workable. The multi-character consistency score is Google’s strongest number, and two or three people in the same frame is exactly where image models have drifted. It still needs testing on your own characters, because the model card itself says consistency is imperfect.
Print and large formats become an option. 4K output at $0.113 makes posters, key art and out-of-home mock-ups possible from the same pipeline, as long as fine text is checked or set by hand.
Disclosure is built in. Every image carries SynthID. That suits clients who want AI work identified, and it is worth saying upfront to clients who do not expect it. Our Responsible AI page covers how we handle it.
Before you switch
Run twenty of your real briefs through both versions and have the person who usually approves images choose blind. Count cost per approved image, retries included, not price per image. If your requests carry many references, check the input cost on a real invoice. There is no rush: as of October 7, Google’s deprecation page lists no shutdown date for Nano Banana 2, even though some coverage reported an October 29 cutoff. The original Nano Banana is the one with a date, March 15, 2027.
Sources
- Nano Banana 2.1 model card (Google DeepMind, October 2026)
- Image generation with Nano Banana (Gemini API documentation, retrieved October 7, 2026)
- Gemini API pricing (Google, retrieved October 7, 2026)
- Gemini API model deprecations (Google, retrieved October 7, 2026)
- Text-to-Image leaderboard (Arena, updated October 6, 2026)
- Image Edit leaderboard (Arena, updated October 6, 2026)
- Google Launches Nano Banana 2.1: Better Than Its Predecessor at Half the Price (Decrypt, October 6, 2026)
- Nano Banana 2.1 Now In Google AI Mode In Search (Search Engine Roundtable, October 6, 2026)
- Google launches Nano Banana 2.1 ahead of the older image API’s earliest shutdown (RuntimeWire, October 6, 2026)