← All articlesAI model reviews

Wan 2.7 review: Alibaba's video model for readable labels and real-looking products

Ryan MitchellSeptember 30, 20268 min read
Wan 2.7 review: Alibaba's video model for readable labels and real-looking products

A hands-on Wan 2.7 review: readable labels, distinct faces, product physics, 15-second clips with sound, Alibaba's API prices and the open-weights catch.

Which model draws a readable label on a box and doesn't give every character the same glossy AI face? If you asked me that tomorrow, I'd say Wan 2.7 before you finished the sentence. That's the short version of this Wan 2.7 review. I'd spent the previous stretch on Hailuo and emotional close-ups, and Wan turned out to be a completely different animal.

My test set was product footage where the text on the packaging had to survive. A brand name on the lid, a logo on the side, a line of copy under it. Most models embarrass themselves here and scribble pseudo-letters. Wan mostly gets it right.

Who makes Wan, and what's this about "thinking"?

Wan comes from Alibaba, specifically its Tongyi Lab. The video side of version 2.7 was announced on April 7, 2026 as a set of four models: text-to-video, image-to-video, reference-to-video and a separate editing model. Alibaba's pitch was all about directing:

"A single prompt can yield a fully realized storyboard enabling users to create director-grade works"

A lot of the launch buzz, including our own notes on it, focused on a Thinking Mode that plans the shot before rendering. Here's something I only noticed while checking the docs for this piece. That thinking_mode switch actually lives in the Wan 2.7 image models. On the video side, the image-to-video API has a prompt_extend option instead. It rewrites your prompt into a fuller one before the render starts, and the docs warn that this adds processing time.

Whatever you call it, the idea is the same. First the model works out the composition and where each object goes. Camera movement gets planned too, and only after that does generation begin. In my runs that planning step paid off on busy scenes with several objects, where fewer things went weird. You pay for it in time. For a simple spinning product shot I'm not sure it's worth the wait.

Three reasons it stays in my rotation

Lettering comes first. Signs, captions, labels on packaging, Wan renders them legibly far more often than most video models I've tried. For ads and e-commerce, where the product has to show its brand name clearly, that's the deciding factor. One caveat. Alibaba's official "12 languages" claim belongs to the Wan 2.7 image model, and I couldn't find a language count for the video models.

Faces are the second reason. You know the effect where everyone in AI footage looks like the same averaged, pretty stranger? Wan gave me people who actually looked different from each other, with different face shapes and eyes. If you're running a campaign with several characters, or a recurring series, that matters more than you'd expect.

Then physics. Wan is good at showing how an object actually works in the frame. The blender actually spins the way a blender should, the iron glides over the shirt without floating above it, and in one of my test clips a drill bites into the board and throws real shavings. For "product in action" footage it's usually my first pick.

The specs you'll actually compare

Clips run from 2 to 15 seconds. Output is 720p or 1080p, according to Alibaba's API docs. Native audio comes out of the same pass with lip sync. You can also feed it your own voice track so the character's mouth and movements follow it. That's a big step up in length from the early Wan releases.

A few other details from the docs:

  • start from a first frame, or pin both the first and the last frame;
  • continue an existing clip;
  • reference-to-video, with up to five distinct characters in one scene per Alibaba's announcement;
  • a dedicated editing model that changes an existing video from a text instruction.

Aspect ratios are the usual 16:9, 9:16 and 1:1 on Flami, plus a few others. That HEX color control you may have read about, where you type your brand color's exact code, is again a feature of the image model. If that's what you need, it's on the Wan 2.7 Image page, and I went through it more closely in the Wan 2.7 Image review.

What it costs

At the source, Alibaba Cloud's Model Studio pricing lists Wan 2.7 text-to-video and image-to-video at $0.10 per second for 720p and $0.15 per second for 1080p in the international region. So a 10-second clip in 1080p comes to $1.50 through the API.

On Flami, Wan 2.7 is included in the regular subscription and paid for in credits. Alongside the flagship you also get Wan 2.6, 2.5, 2.2 Fast and 2.2. My logic is simple enough. 2.7 gives you the best results but takes longer and costs more, while the older versions are quicker and cheaper for bulk jobs.

The open-weights catch

Wan built its reputation as an open model. People downloaded the weights and ran it on their own GPUs. That's no longer the case for the newer releases. Alibaba's Wan-AI page on Hugging Face still stops at the 2.2 family from last summer, and 2.5, 2.6, 2.7 and the new 3.0 are available through the API only.

If you use Wan through a service, that changes nothing for you. If you were planning to host 2.7 on your own hardware, you're out of luck, and 2.2 is the newest version you can still run locally.

Wan 3.0 is out, so does 2.7 still matter?

Alibaba released Wan 3.0 in August. It makes clips up to 30 seconds and accepts a pile of new inputs, even documents. It's also near the top of the blind rankings. In the September 21, 2026 update of arena.ai's text-to-video board, wan3.0 is 6th with 1476, while wan2.7-t2v sits 20th at 1337.

I haven't tested 3.0 properly yet, so I'll stick to what I know. 2.7 is still a solid pick for product shots with text on them, and at list prices it's a bit cheaper per second than 3.0 in 1080p. I'd probably revisit this once I've put the two side by side on the same products.

How I split jobs between Wan and the rest

To keep things straight, this is roughly how I assign work:

  • text, logos and packaging in the frame: Wan 2.7;
  • a product doing its job, with believable physics: Wan 2.7;
  • premium light and a film look: Veo 3.1;
  • emotion in a close-up: Hailuo;
  • fast motion and camera fly-arounds: Kling 3.0.

Asking whether Wan or Veo is "better" in general doesn't get you far. It depends what's failing on your shot: dull light or a blurry label.

Should you pick Wan 2.7?

I reach for 2.7 whenever the box has real text or a logo on it, when the scene needs distinct faces instead of AI clones, or when the product has to actually work on screen. The 15-second single pass helps too.

I skip it when someone wants to run the model on their own server, since the 2.7 weights were never published, or when the whole point is cinematic light and mood. That one goes to Veo.

What's next

Seedance 2.0 is next on my list, then HappyHorse 1.0 and Grok Imagine. I also want to show a combo I've been using, with a base clip in Wan and finishing in Runway. If you haven't used Runway as an editor before, the Runway Aleph review explains the idea.

Load a labelled box into Wan 2.7 on Flami and zoom into the text to see how it holds up.

Things people ask about Wan 2.7

  • What is Wan 2.7? Alibaba's video model family announced in April 2026. It makes clips from 2 to 15 seconds with native sound, and it's especially good at text in the frame and at distinct faces.
  • Does Wan 2.7 have a thinking mode for video? Alibaba documents the thinking_mode switch for the Wan 2.7 image models. The video API has prompt_extend instead. It expands the prompt before rendering and takes longer.
  • What length and resolution? Anywhere from 2 to 15 seconds. Resolution is 720p or 1080p, with native audio and lip sync.
  • Can I run Wan 2.7 locally? No. Its weights aren't published, and 2.2 is the newest Wan with open weights.
  • How much does it cost through Alibaba's API? $0.10 per second at 720p and $0.15 at 1080p in the international region. On Flami you pay in credits.
  • How is it different from Veo and Kling? Wan doesn't compete with either on light or top-end motion, its edge is text staying readable and objects behaving correctly, and unlike older Wan releases, 2.7's weights aren't public, so self-hosting is off the table.

Last checked on September 30, 2026, against Alibaba Cloud's docs and the arena.ai video board.

Sources

  1. Alibaba Cloud, Alibaba Unveils Wan2.7-Video
  2. Alibaba Cloud, Alibaba Unveils Wan2.7 image generation
  3. Alibaba Cloud Model Studio, Wan 2.7 image-to-video API reference
  4. Alibaba Cloud Model Studio, Wan 2.7 image generation and editing API reference
  5. Alibaba Cloud Model Studio, model pricing
  6. Alibaba Cloud, Wan3.0 launch
  7. Hugging Face, Wan-AI models
  8. arena.ai, Text-to-Video Leaderboard
  9. Flami, Wan 2.7 launch notes
  10. Flami, Wan 2.7 page

About the author

Ryan Mitchell

Reviewer at Flami

Read next

Get 15 credits for free

Use them to generate images and videos:

≈ 11 × Nano Banana 2, ≈ 12 × GPT Image 2, ≈ 1 × Seedance 2

Get free credits