HappyHorse 1.0 Review: Where This Video Model Nails Product Shoots (and Where It Doesn't)

HappyHorse 1.0 on product shoots: how it handles physics, reflections, and fine detail in motion, what it nails, and where it still falls short.
I'm Alex, a product engineer at Flami. Part of my job is running video models through product-shoot tests and tracking where they hold up and where they fall apart. In early April, a model with the odd name HappyHorse showed up on the Artificial Analysis arena with zero branding and just took the top spot. Nobody knew whose it was at first. A couple of days later, Alibaba admitted it was theirs. My first reaction was skepticism: anonymous models pop up on leaderboards all the time and quietly vanish a week later. This one didn't vanish. Once I started running it against real product footage, I understood what the fuss was about.
Short version: when the shot centers on a product and you need detail that doesn't warp or smear on motion, HappyHorse is now one of my top three picks.
Where This Horse Came From
On April 7, HappyHorse-1.0 showed up in the Artificial Analysis video arena with no listed creator and immediately passed the previous leader, Seedance 2.0. On April 10, Alibaba confirmed to CNBC that HappyHorse is its project, built by the ATH team inside Taotian Group, and that the model is still under active development. The team is led by Zhang Di, former VP at Kuaishou and the person behind Kling. So this isn't a side project from a few grad students, it's a team that already shipped one leading video model. I covered the announcement and benchmark standings in a separate piece; here I only care about what it does on product work.
Under the hood, going by what the team has published, it's a single transformer with 15 billion parameters that generates video and audio in one pass, with no separate post-processing stage. Most competitors still stitch the result together across a pipeline. This one runs on a single engine. The team claims support for seven languages with lip sync, and a generation speed around 38 seconds per 1080p clip on one H100.
Why I Put It Through a Product Shoot Test
Benchmarks are one thing. What I actually need is simpler: a product on screen that looks real and doesn't come apart the moment it moves.
Motion stability. I ran the prompt "a person spins a hula hoop and squats." On most models, the hoop either flies out of frame or turns into a smear. HappyHorse kept the hoop's geometry and the body's proportions intact through the whole squat. No drift either: no slow warping where an object quietly floats out of place over time.
Reflections. This one surprised me. A mug on a glossy table, the reflection on the surface moves in sync with the object and keeps its proportions and lighting. Mirrors, chrome, water, HappyHorse's reflection handling actually tracks the geometry instead of smearing a blurry patch. For fashion shots or premium products, where half the frame is reflective surface, that's a real advantage.
Liquids and texture. Coffee foam, latte art, droplets, smoke. Liquid follows surface physics instead of jittering frame to frame. I ran a clip of a drink being poured, and the spread looked convincing, no stutter.
What the Leaderboard Numbers Actually Say
Here's the part marketing likes to gloss over.
In text-to-video without audio, HappyHorse-1.0 holds the top spot on Artificial Analysis with an Elo around 1357, well ahead of Seedance 2.0. Image-to-video without audio shows the same pattern, a clear lead. Turn audio on, though, and the gap closes: in text-to-video with audio, HappyHorse and Seedance 2.0 from ByteDance are nearly tied, a difference of one or two Elo points, basically noise. In image-to-video with audio, Seedance edges ahead in some runs.
"HappyHorse-1.0 currently leads the Artificial Analysis Text to Video Arena (without audio) with an Elo score of 1357."
Source: Artificial Analysis, Text to Video Leaderboard
So "leaderboard leader" is true, with a caveat: it's unambiguously ahead on visuals without audio. With audio on, it's a tie with Seedance, not a blowout. For product work that's not a problem, most product videos go through editing with separate voiceover or music anyway, so visual quality without audio is exactly what I care about first.
HappyHorse Against the Rest, Rough Breakdown
Here's roughly how I pick a model depending on the job:
- fine product detail, motion stability: HappyHorse 1.0
- reflections in mirrors, chrome, water: HappyHorse 1.0
- close-up emotion, facial expression: Hailuo
- premium look, cinematic lighting: Veo 3.1
- dynamic camera movement: Kling 3.0
- object physics in hand: Wan 2.7
- phone-shot UGC look, audio baked into the frame: Seedance 2.0
None of this means HappyHorse is weaker overall. Each model has its strength. HappyHorse's is geometry and surfaces holding together under motion.
Up to Nine Characters in One Scene
There's a feature I underrated at first. HappyHorse has a reference-based mode: feed it up to nine reference images and use them as distinct characters via tags like "character1," "character2." Each one keeps its appearance consistent through the scene.
For product work, that's not a gimmick. A group shot with several products interacting, or a scene with two models and a product passed between them, used to mean stitching separate generations together in editing. Here you can try it in one pass. I've only tested up to three references so far, not the full nine, so I won't vouch for stability at the full count. But even at three, the consistency held.
Length, Formats, Audio
Spec sheet: clips run 3 to 15 seconds, default 5. Resolution is 720p or 1080p. Aspect ratios cover everything you'd need, 16:9 for horizontal, 9:16 for Reels and TikTok, 1:1 for Instagram, plus 4:3 and 3:4.
Audio generation is optional, you choose per project. For product videos I usually skip it and add my own music and voiceover in editing, it gives more control. Input options are text, up to nine reference images, and optionally an existing video for editing.
Speaking of editing, there's a mode where you give a finished clip a text instruction, swap an object, change the style, make a local edit without regenerating the whole thing from scratch. It accepts video from 3 to 60 seconds. I've only tried this lightly so far, looks convenient, but how it holds up on a real product clip with fine detail still needs more testing on my end.
Now the Honest Part, Where It Falls Short
I promised no cheerleading, so here are the weak points.
First, and this matters most: at launch, the model was described as open but the weights hadn't actually been released, GitHub and Model Hub links were still marked "coming soon." Alibaba itself says the project is still in development. I'd hold off treating it as a finished, stable product since it's brand new.
Second, independent testing notes that scene continuity degrades on longer clips: the farther into a generation, the higher the chance some movement logic breaks. I caught this too on eight-second clips, things occasionally wobble near the end. For a 5-6 second product shot that barely matters, but if you're planning a longer continuous sequence, keep it in mind.
Third, a subtler one. Reviewers point out that HappyHorse has a compressed color space, which limits flexibility in color grading, less headroom in highlights, saturated reds and cyans clip sooner. If you're doing serious grading work afterward in Resolve, the footage gives you less room than real camera footage would. Our motion designer flagged this to me directly; I don't work deep in color grading myself, so I'm passing it along secondhand, and it's possible a later version fixes it.
And on audio, as I said, it's not the decisive win it is on visuals. The audio is usable, just not ahead of Seedance.
Pricing and Where to Get It
Through third-party platforms, billing runs per second: $0.14 per second at 720p, $0.28 per second at 1080p, no minimums, no subscription required. A ten-second 1080p clip runs around $2.80.
In Flami, HappyHorse 1.0 is part of the shared subscription, billed through credits. You don't need to sign up on a separate platform or deal with per-second billing just to try one model. It sits in the same workspace as Veo, Kling, Hailuo, and the rest, so you switch models depending on the job instead of juggling accounts.
Checklist: Should You Use HappyHorse 1.0
- ✓ Product shoots where detail and motion stability matter
- ✓ You need convincing reflections in mirrors, chrome, glass, or water
- ✓ Complex physics: sports, acrobatics, interaction with objects
- ✓ Multiple distinct characters in one scene from reference images
- ✗ The shot relies on close-up facial emotion, Hailuo is stronger there
- ✗ You need maximum flexibility for color grading, the color space is compressed
- ✗ A long continuous scene in one pass, continuity can break near the end
What's Next
I want to run a direct comparison between HappyHorse and Seedance 2.0 on the same product prompts with audio on, since they're tied on the leaderboard but the difference might show up more clearly side by side. Also in the queue: a look at Wan 2.7 and a note on Grok Imagine. Different articles, though.
You can try HappyHorse 1.0 in Flami, free starter credits included.
Frequently Asked Questions
What is HappyHorse 1.0? A video generation model from Alibaba's ATH team, announced in April 2026. It produces clips from 3 to 15 seconds in 720p or 1080p, generates video and audio in a single pass, and supports four modes (text, image, reference-based, and editing an existing video) across seven languages.
Is HappyHorse really a leaderboard leader? Yes, with a caveat: HappyHorse holds the top spot on Artificial Analysis for text-to-video and image-to-video without audio, while with audio on it's tied with Seedance 2.0, and in image-to-video with audio Seedance is sometimes ahead.
How does HappyHorse compare to Seedance, Kling, and Veo? Its strengths are stable physics on complex motion, accurate reflections in mirrors and chrome, correct liquid simulation, and strong product detail. On visual quality without audio, it currently ranks above Seedance 2.0, Kling 3.0, and Veo 3.1.
How many characters can appear in one scene? Up to nine. In reference mode, you supply between one and nine reference images and tag them to characters using labels like "character1," "character2," and so on. The model keeps each one's appearance consistent through the scene.
What are the weak spots? The weights aren't publicly released yet, the model is new and still under development. Longer clips lose scene continuity toward the end. The color space is compressed, giving less room for grading. Audio is usable but not ahead of Seedance.
Does it support multiple languages? Yes, you write scene and style directly in the prompt. The team claims multi-language support with lip sync across seven languages.
How do I access HappyHorse 1.0? Directly through third-party platforms, you'd pay per second with no bundled access to other models. Through Flami, it's included in the shared subscription with credit-based billing alongside the other models, and new accounts start with free credits.
Sources
- CNBC: Alibaba revealed as creator of AI video generation model HappyHorse-1.0
- Artificial Analysis: Text to Video Leaderboard
- Artificial Analysis: Image to Video Leaderboard
- South China Morning Post: Alibaba's HappyHorse tops Seedance
- Flami: HappyHorse 1.0 product page
About the author
Ryan Mitchell
Reviewer at Flami