Veo 3.1 review: what Google's video model does well in 2026, and where it still slips

A hands-on Veo 3.1 review: native sound, Extend for longer clips, one character across scenes, weak spots, Google's API prices and when it's worth it.
Is a video model still worth paying for once it's slipped out of the top ten? That's the question I kept asking myself while writing this Veo 3.1 review. I test video models for Flami before they reach users, which mostly means running them by hand and then counting where they fib. Veo 3.1 got two straight weeks of that from me, first on my own little projects and then on a batch of product videos. Below is what it can actually do, where it still messes up, and why you'd bother with it if you already have Kling or Runway.
Two clips from that test first, so you've got something to look at.
What Veo 3.1 is, in plain terms
It's the video model from Google DeepMind. Version 3.1 came out on October 15, 2025, announced on the Google Developers Blog and covered the same day by TechCrunch. Since then Google has kept patching it rather than replacing it. A January update improved character consistency and added native vertical video plus 1080p and 4K upscaling, and in March a cheaper Veo 3.1 Lite tier showed up for developers. As of today, DeepMind's model page still lists 3.1 as the current Veo. There's no Veo 4, whatever some blogs claim.
Where does it stand against the competition? Lower than it did last fall. On arena.ai's video board (September 21, 2026 update) the best Veo 3.1 entry, the audio version, is 12th with 1364. Newer models sit above it, including one of Google's own.
The specs, per the Gemini API documentation:
- clip length: 4, 6, 8 seconds;
- resolution: 720p by default; 1080p or 4K only for 8-second clips;
- frame rate: a fixed 24 fps;
- aspect ratios: 16:9 and 9:16;
- reference images: up to three.
For most jobs 8 seconds is plenty. When it isn't, there's Extend.
Google names four headline features for this version, so I'll take them one at a time and say what held up in my tests.
Native sound is the big one
Video first, then the audio written or generated separately, then a mix pass to tie them together. That's the old workflow. Veo 3.1 does all of it in one generation.
And it's synced to the scene, which is the point. Water pouring in the frame means you hear it pour. A person talking gets lip sync. I tested speech mostly in English, where lip sync was good. Google's docs say English is the only language they've fully evaluated. The exact wording is that other languages "may work but results can vary," and that tracked with what I'd expect for anything outside English. For anything important in another language I'd either write the dialogue in English or skip speech and keep only ambient sound.
Ambient sound is close to flawless. Café chatter, footsteps on pavement, an espresso machine hissing, the rustle of fabric. The model adds all of that on its own from the context. For simple clips under 8 seconds you really don't need a sound designer anymore.
What didn't match the promises was stereo. I expected a proper spatial mix. What I got was mono or fake stereo. Maybe Quality mode fixes that, I mostly ran Fast, so I might be wrong here.
Extend, or how to get past 8 seconds without visible seams
A 20 or 30-second ad meant generating several pieces and stitching them in an editor, and you could always see the joins. The light shifts, the character's pose jumps, sometimes the skin tone changes, and viewers' eyes snag on exactly those moments.
Extend continues a finished clip so the next segment starts precisely on the frame where the previous one ended. Lighting carries over, so do the camera angle and the color. The API adds 7 seconds per extension, up to 20 times, for a ceiling of 148 seconds, though only at 720p.
In practice I could chain 3 or 4 pieces before things went wrong. After that the model slowly loses the character's face and details of the outfit start to change. So 16 to 24 seconds is the range I'd actually trust, even if the spec sheet allows much more. On Flami, Extend is a separate step you run after the base clip, right from the Veo 3.1 page.
One character across several scenes
This is a big step up from Veo 3. One reference image used to be the limit. Veo 3.1 takes up to three, so you can lock in a person, an outfit, a whole look, and carry it through the clip. Google's January update pushed this further and calls it identity consistency.
The model keeps the face and the clothes consistent throughout the scene. Even the way someone moves. If you're making a campaign with one recurring character, or anything episodic, that's critical.
Close-ups worked really well for me. On wide shots the character sometimes drifts in small details, especially when my reference images were taken from different angles. My advice is to upload references of one type. Either all portraits or all full-length, just don't mix them.
How is "narrative control" different from a normal prompt?
Subtle point, but it matters. In the Flow announcement Google talks about more narrative control, and in practice that means you can pin specific events to specific moments inside the 8 seconds.
A normal prompt describes the overall idea. What happens, the style, the mood. With Veo 3.1 you can be literal: say the hero turns his head two seconds in, someone else steps into frame a beat later, and by the end of the clip they're looking at each other. The model sticks to those beats.
I didn't get this right away. My first tries were just long prose prompts, and they didn't work. You have to spell out the timings, and then it listens. Rough example. Instead of "a barista makes coffee and smiles at the customer," I'd write "The first couple of seconds, beans into the grinder. Then a longer stretch pouring the milk. Right before the cut, she smiles."
Where it falls short
I'm not going to pretend it's perfect.
Hands doing complicated things are a weak spot. When someone is arranging things or typing, the fingers sometimes smear. Kling 3.0 handles that better, to be fair.
Readable text in the frame is another. Want a legible label with the brand name on a bottle? Not happening. Veo writes pseudo-letters. I'd make that shot in an image model like Ideogram or GPT Image and then feed it in as a reference.
Then cost. Google's Gemini API price list puts standard Veo 3.1 at $0.40 per second of video at 720p or 1080p and $0.60 at 4K, while Fast runs $0.10 to $0.12 per second at those lower resolutions. So a single 8-second 1080p clip in the top mode costs $3.20 at API rates. If I had a series of 200 SKUs to do, I'd run it on Kling and save Veo for the hero products where a premium look actually pays off.
Speed too. Fast takes 1 to 2 minutes, Quality 3 to 5. Averaged across our dashboard it's around 2 to 4 minutes a clip. Fifty clips in one go is about two hours of waiting.
Jobs I'd give it
After two weeks this is how I split things for myself.
Hero product pages. When you have 5 or 10 top products and each one carries real budget, Veo does it best. The picture reads as cinematic, the light stays natural, and textures actually look real instead of rendered.
Lifestyle and mood-driven categories like fragrance, premium skincare, higher-end clothing. Here emotion and setting matter more than action, and Veo handles that nicely.
Explainers and brand content. For tutorials and brand content on YouTube, Veo gives you a "shot on a real camera" feel without that floaty AI motion.
What I wouldn't do on Veo is bulk product clips for a marketplace. Too expensive for that, and too slow. Kling or Hailuo handle it for noticeably less money.
Using it through Flami
On Flami, Veo 3.1 is part of the general subscription, so there's no separate Google account to set up and no Google Cloud project to configure. That matters more than it sounds like it would. Going direct means an API key plus a billing account plus working through code. The other route is Google's Flow app with its own plan.
There are two modes to choose from. I used both a lot. Fast is for drafts and prompt tests. Quality is for when you already know what you want and need the final version.
Prompts in other languages will usually get understood. For complex scenes with dialogue, English is safer, as I mentioned above.
What it costs on Flami
We don't list a dollar price for international accounts yet, so I'll keep this simple. You pay in credits. The cost depends on mode and length, resolution too. One Extend costs about as much as another generation. Current numbers are on the Veo 3.1 page.
Quick check: is Veo 3.1 right for you?
Good fit if you need:
- a premium, cinematic look;
- sound that's synced to the picture;
- ads for luxury or beauty products;
- longer clips, 15 to 30 seconds, with one recurring character.
Probably not the right tool if:
- you're making a series of 50+ SKUs, it'll get pricey fast, try Kling;
- you need readable text in the frame, use a different model for that shot;
- the whole clip is about hands handling objects, where Kling will do better.
Where to go from here
If you're choosing between models, I put Veo up against Kling and Runway on the same product shots in a three-way comparison. Kling also got its own write-up if that's the one you're leaning toward.
Or just try it yourself. Sign up on Flami and burn a few Fast drafts on the free credits before you commit to Quality.
FAQ
*What is Veo 3.1?* Google DeepMind's video generation model, released October 15, 2025. It makes realistic clips with sound, keeps a character consistent across scenes and can continue a clip through Extend.
*How is Veo 3.1 different from Veo 3?* Three additions stand out: first and last frame control, several source images per request, and Extend for clips longer than 8 seconds.
*Veo 3.1 vs Sora 2, which is better?* On arena.ai's video board they're practically tied. Sora 2 Pro is 11th with 1368 and the Veo 3.1 audio version is 12th with 1364, which is well inside the margin of error. My own impression is that Veo follows the prompt more literally.
*Does it work in languages other than English?* Prompts usually get understood. Google has only fully evaluated English, though, so dialogue in other languages can be hit or miss.
*How many seconds per generation?* Up to 8. Through Extend the API allows up to 148, but in my tests it stayed coherent to around 24 or 32.
*Can I upload my own photos?* Yes, up to three reference images per request, which is how you keep a character or a style consistent.
*What resolution does it support?* 720p by default. 1080p or 4K only for 8-second clips.
Last checked against Google's docs on September 30, 2026.
Sources
- Google DeepMind, Veo
- Google Developers Blog, Introducing Veo 3.1 and new creative capabilities in the Gemini API
- Google, Introducing Veo 3.1 and advanced capabilities in Flow
- Google, Veo 3.1 Ingredients to Video update
- Google AI for Developers, Veo 3.1 documentation
- Google AI for Developers, Gemini API pricing
- TechCrunch, Google releases Veo 3.1, adds it to Flow video editor
- arena.ai, Text-to-Video Arena
- Flami, Veo 3.1 page
About the author
Ryan Mitchell
Reviewer at Flami
Read next
Kling 3.0 review: two weeks of product clips, multi-shot scenes and a few misfires
Runway Aleph review: an AI video editor, not another text-to-video model
Wan 2.7 review: Alibaba's video model for readable labels and real-looking products
Grok Imagine review: the video model I use when I haven't decided yet