Grok Imagine review: the video model I use when I haven't decided yet

A hands-on Grok Imagine review: clips with sound in seconds, 1080p in Video 1.5, reference images, xAI's API prices, and the jobs where slower models still win.
By the time a heavier model finishes its first take, Grok Imagine has usually already run through three or four passes on the same prompt, quick enough that I lose track of takes before my coffee even cools. That speed gap is the whole reason this review exists. Grok has quietly become the model I keep open in a spare tab, mostly because so much of my day goes to not yet knowing which offer will land.
If you want xAI's official specs, I pulled them together in a separate breakdown. This piece is about how the model behaves on real work.
Fast enough to change how you work
On Flami a standard clip usually lands in 5 to 20 seconds, and heavier scenes take up to about 30. Premium models in my queue sit at two to six minutes. That gap sounds like a convenience. In practice it changes the whole approach.
Five minutes of staring at a progress bar per miss is why I used to polish every prompt before submitting it. On Grok Imagine I stopped doing that and just try things. Rough idea, look, adjust the prompt, run it again. I can go through that loop twenty times in the window a heavy model spends on one clip.
Where does that pay off most? A/B tests, mainly, when I need to see ten versions of one offer instead of guessing which one will land. Social content leans on it too, since posting at feed speed doesn't wait for a four-minute render.
Quality at that speed
I expected it to look cheap. It doesn't, mostly.
Back on June 16, 2026, xAI shipped Grok Imagine Video 1.5 and called it its best image-to-video model so far. Sound effects come out of the same pass as the picture, along with ambience and speech. On July 31 an update added native 1080p plus text-to-video, so you no longer need a starting image. On Flami clips run at 24 frames per second in 16:9 or 9:16, or square if you're making something for an Instagram feed. I almost never have to build an audio track separately for a draft.
That same July update brought reference images, up to seven per generation. xAI's wording is that each reference "locks one thing in place," and the examples they give include a face or a product. For product clips that's handy. Upload a photo of a mascot or a model and it carries through the shot without me writing a paragraph describing their face. There's a catch in the video generation docs, though. Reference-to-video is capped at 720p, and only plain text-to-video and image-to-video get the full 1080p.
Camera moves go in plain English. Dolly zoom, a slow push past the product, a rack focus. You don't need any special markup, and it usually listens. The docs also describe pinning a last frame, or up to four keyframes at set timestamps, which I've only played with a little.
Blind voters seem to agree it's no longer the budget pick. In my notes on the September update of the arena.ai image-to-video board, Video 1.5 at 720p came in eighth with 1456.
Where it gives way
Something has to give when a model is this quick.
Length, for starters. xAI's API allows clips of 1 to 15 seconds, and on Flami you pick from 6 to 15. For a longer story I'd need to stitch, and that's where speed stops helping.
For a premium shot with perfect light and a busy scene, I'd rather build it on Veo 3.1. Its footage simply looks more expensive, and the Veo 3.1 review gets into why. Tricky physics is another soft spot. Reflections, liquids, anything that splashes. For that I'd try HappyHorse 1.0 first. And for a close-up talking head where every syllable has to line up with the mouth, Grok honestly isn't my first choice either, since I've seen the articulation drift on longer lines. I haven't put 1.5 through enough talking-head shots to call that fixed, so take the critique with a grain of salt.
Price per second isn't the problem. xAI's API pricing for Video 1.5 is $0.08 per second at 480p, $0.14 at 720p and $0.25 at 1080p, so a 10-second 1080p clip costs $2.50 at list price. On Flami it's billed in credits like everything else.
Grok first, heavy model second
For me Grok belongs to the "still thinking" stage. When I don't know yet which offer will hit, I don't polish one clip. I rough out ten and watch the reaction.
The usual routine goes like this. At the start of a campaign I run a batch of creatives through Grok and see which offer and which framing people stop for. Only the winner gets rebuilt properly on something heavier, and only if it earns it. Nobody burns half a day on a clip that gets two seconds of attention in the feed, and the credits stay mostly intact. I wrote about other places to save money on video in the piece on fast, cheap video models.
Using it without an xAI account
On Flami, Grok Imagine comes with the regular subscription, so there's no separate xAI signup or plan to manage. You can start from a text prompt or a source image. If typing feels slow, you can dictate the prompt by voice, which comes in handy more often than I expected.
I still get a small kick out of roughing out fifteen versions of one offer in the time a heavy model spends on a single clip. It isn't the prettiest model in the lineup, and I still open it more than any other, probably because most of my days are spent not knowing yet.
Grok Imagine questions
What is Grok Imagine?
A video generation model from xAI. The current version, Grok Imagine Video 1.5, went generally available on June 16, 2026. It makes clips with synchronized audio, and since July 31 it supports native 1080p for text-to-video and image-to-video.
How long can a Grok Imagine clip be?
The xAI API accepts 1 to 15 seconds per clip. On Flami you choose between 6 and 15 seconds.
How much does Grok Imagine cost through the API?
Per xAI's price list, Video 1.5 costs $0.08 per second at 480p, $0.14 at 720p and $0.25 at 1080p. Using an image or video as input is charged separately.
Can it keep the same character across clips?
Yes. Since the July 31 update you can attach up to seven reference images, for example a face or a product. Reference-to-video output is capped at 720p.
Do I need an xAI subscription to use it?
Not on Flami. It's included in the regular subscription alongside the other video models.
Last checked on September 30, 2026. Specs and prices come from xAI's announcements plus its API docs.
Sources
About the author
Ryan Mitchell
Reviewer at Flami