Happy Horse 1.0 is Alibaba's chart-topping generator that produces 1080p video and its synchronized audio in a single pass — dialogue, Foley, and ambient sound included, no dubbing step.
A unified single-stream Transformer emits video frames and their matching soundtrack jointly in one forward pass, so lip movement, dialogue, and ambience are aligned by design rather than stitched afterward.
Debuting anonymously, it topped the Artificial Analysis Video Arena in both text-to-video and image-to-video on blind human preference votes, overtaking previous leaders.
Generates emotionally expressive, lip-synced dialogue across seven languages including English, Mandarin, Cantonese, Japanese, Korean, German, and French.
Alibaba announced the model would be released under Apache-2.0, positioning it as a state-of-the-art open-source alternative to closed commercial video models.
Happy Horse 1.0 (stylized HappyHorse-1.0, Chinese 快乐小马) is an AI video generation model from Alibaba that made an unusual entrance: it appeared at the top of the Artificial Analysis Video Arena leaderboard in April 2026 without any company attached to it, and within days Alibaba confirmed it was behind the model. Built by an ATH-AI Innovation team led by Zhang Di — a former Kuaishou VP and technical architect of Kling AI — it quickly became one of the most talked-about releases of the year, claiming the top ranking in both text-to-video and image-to-video categories on blind human-preference voting.
What sets Happy Horse 1.0 apart technically is native joint audio-video generation. Rather than producing a silent clip and dubbing sound in a separate step, its roughly 15-billion-parameter single-stream self-attention Transformer generates the video and its synchronized soundtrack — dialogue, Foley, and ambient sound — together in one forward pass. Combined with native 1080p output, multi-shot sequences, and lip-synced speech across seven languages, this makes it aimed squarely at end-to-end, ready-to-watch video rather than raw visuals that need a post-production pipeline.
The model's defining feature is generating picture and sound together in a single pass rather than as two stages. Because audio is produced alongside the frames, lip movement, spoken dialogue, sound effects, and background ambience stay coherent without a separate dubbing or Foley step.
Reports describe a roughly 15-billion-parameter single-stream self-attention Transformer that handles both modalities in one architecture, rather than bolting an audio module onto a separate video generator. This unified design is credited for the tight audiovisual synchronization.
Happy Horse 1.0 generates in HD 1080p natively rather than upscaling from a lower base resolution. It also supports multi-shot sequences, allowing a single generation to move between camera angles within one clip.
The model produces lip-synced, emotionally expressive dialogue across seven languages — English, Mandarin, Cantonese, Japanese, Korean, German, and French. This makes it well-suited to talking-character and localized content where mouth movement must match the spoken track.
It accepts both text prompts and still images as input. Image-to-video lets creators animate an existing frame or character design, while text-to-video builds a scene from a description — both with the same native-audio pipeline.
On the Artificial Analysis Video Arena, which ranks models by blind human preference, Happy Horse 1.0 reached the top position in both text-to-video and image-to-video, ahead of previously leading models. It is one of the highest-rated video models on that leaderboard to date.
Alibaba announced plans to release Happy Horse 1.0 under the permissive Apache-2.0 license, framing it as an open-source, state-of-the-art option. This positions it differently from closed commercial video models that are API-only.
| Name | Happy Horse 1.0 (HappyHorse-1.0) |
|---|---|
| Developer | Alibaba, ATH-AI Innovation team |
| Parameters | ~15 billion |
| Architecture | Single-stream self-attention Transformer |
| License (announced) | Apache-2.0 |
| Inputs | Text prompt, image |
|---|---|
| Output resolution | 1080p HD (native) |
| Audio | Native synchronized audio (dialogue, Foley, ambience) |
| Multi-shot | Supported within a single clip |
| Lip-sync languages | English, Mandarin, Cantonese, Japanese, Korean, German, French |
| Leaderboard | Artificial Analysis Video Arena |
|---|---|
| Text-to-video rank | #1 (blind human preference) |
| Image-to-video rank | #1 (blind human preference) |
| Debut | April 2026 (anonymous, later attributed to Alibaba) |
Because lip-sync and speech are generated with the video, the model suits avatars, explainers, and character scenes where mouth movement must match spoken lines — across seven languages without a separate voiceover pass.
Native audio means short-form clips arrive with ambience and sound effects already attached, cutting the editing loop for creators publishing quickly to social platforms.
Multilingual lip-sync makes it practical to produce the same scene in several languages with matching mouth movement, useful for global marketing and education content.
Image-to-video lets designers and studios bring a single frame, product shot, or character illustration to life with motion and a synchronized soundtrack.
Multi-shot 1080p generation with sound helps teams sketch out sequences, mood, and pacing early — a fuller preview than silent draft renders.
Happy Horse 1.0 overtook ByteDance's Seedance 2.0 at the top of the Artificial Analysis Video Arena. Here's how the two leading models line up on the headline capabilities.
| Feature | Seedance 2.0 | Happy Horse 1.0NEW |
|---|---|---|
| Video Arena rank (T2V, at HH debut) | Previous leader | #1 |
| Native synchronized audio | Not the core focus | Yes, joint audio-video in one pass |
| Multilingual lip-sync | Limited | 7 languages |
| Announced open-source license | Closed / commercial | Apache-2.0 (announced) |
| Developer | ByteDance | Alibaba |
Alibaba announced an Apache-2.0 release, but at the time of the sources reviewed the full model weights had not yet been publicly downloadable — the repository and model pages existed without complete artifacts. Availability may lag the announcement.
Reported Elo scores and exact rankings differ between leaderboard captures and categories. Treat specific point values as approximate; the consistent takeaway is a top-tier, #1-class placement rather than one fixed number.
Like most current video models, generations are short clips rather than long-form video, and complex multi-shot prompts can still produce continuity or artifact issues. Plan for iteration on demanding scenes.
Much of the early availability came through API partners and aggregators rather than a single official endpoint, so pricing, quotas, and feature parity can differ by provider until first-party access matures.
We're adding Happy Horse 1.0 as platform access opens — explore SharkFoto's available AI tools today and check back for availability.
Try Happy Horse 1.0 now