Alibaba · Video and audio, generated together.

Happy Horse 1.0

The #1 AI Video Model With Sound Built In

Happy Horse 1.0 is Alibaba's chart-topping generator that produces 1080p video and its synchronized audio in a single pass — dialogue, Foley, and ambient sound included, no dubbing step.

15B
Parameters
1080p
Native output
7
Lip-sync languages
#1
Video Arena rank
What's new

The #1 AI Video Model With Sound Built In

01

Native audio-video

A unified single-stream Transformer emits video frames and their matching soundtrack jointly in one forward pass, so lip movement, dialogue, and ambience are aligned by design rather than stitched afterward.

02

Ranked #1 on the arena

Debuting anonymously, it topped the Artificial Analysis Video Arena in both text-to-video and image-to-video on blind human preference votes, overtaking previous leaders.

03

Multilingual lip-sync

Generates emotionally expressive, lip-synced dialogue across seven languages including English, Mandarin, Cantonese, Japanese, Korean, German, and French.

04

Open-source ambitions

Alibaba announced the model would be released under Apache-2.0, positioning it as a state-of-the-art open-source alternative to closed commercial video models.

Overview

Model at a glance

Happy Horse 1.0 (stylized HappyHorse-1.0, Chinese 快乐小马) is an AI video generation model from Alibaba that made an unusual entrance: it appeared at the top of the Artificial Analysis Video Arena leaderboard in April 2026 without any company attached to it, and within days Alibaba confirmed it was behind the model. Built by an ATH-AI Innovation team led by Zhang Di — a former Kuaishou VP and technical architect of Kling AI — it quickly became one of the most talked-about releases of the year, claiming the top ranking in both text-to-video and image-to-video categories on blind human-preference voting.

What sets Happy Horse 1.0 apart technically is native joint audio-video generation. Rather than producing a silent clip and dubbing sound in a separate step, its roughly 15-billion-parameter single-stream self-attention Transformer generates the video and its synchronized soundtrack — dialogue, Foley, and ambient sound — together in one forward pass. Combined with native 1080p output, multi-shot sequences, and lip-synced speech across seven languages, this makes it aimed squarely at end-to-end, ready-to-watch video rather than raw visuals that need a post-production pipeline.

Model
Happy Horse 1.0 (HappyHorse-1.0)
Developer
Alibaba (ATH-AI Innovation team)
Team lead
Zhang Di (ex-Kuaishou / Kling)
Architecture
~15B single-stream Transformer
Inputs
Text and image
Output
1080p video with native audio
Languages (lip-sync)
7 (EN, ZH, Cantonese, JA, KO, DE, FR)
Leaderboard debut
April 2026, Artificial Analysis Video Arena
Capabilities

Key features

Joint audio-video generation

The model's defining feature is generating picture and sound together in a single pass rather than as two stages. Because audio is produced alongside the frames, lip movement, spoken dialogue, sound effects, and background ambience stay coherent without a separate dubbing or Foley step.

Unified single-stream Transformer

Reports describe a roughly 15-billion-parameter single-stream self-attention Transformer that handles both modalities in one architecture, rather than bolting an audio module onto a separate video generator. This unified design is credited for the tight audiovisual synchronization.

Native 1080p output

Happy Horse 1.0 generates in HD 1080p natively rather than upscaling from a lower base resolution. It also supports multi-shot sequences, allowing a single generation to move between camera angles within one clip.

Multilingual lip-sync

The model produces lip-synced, emotionally expressive dialogue across seven languages — English, Mandarin, Cantonese, Japanese, Korean, German, and French. This makes it well-suited to talking-character and localized content where mouth movement must match the spoken track.

Text-to-video and image-to-video

It accepts both text prompts and still images as input. Image-to-video lets creators animate an existing frame or character design, while text-to-video builds a scene from a description — both with the same native-audio pipeline.

Benchmark-topping quality

On the Artificial Analysis Video Arena, which ranks models by blind human preference, Happy Horse 1.0 reached the top position in both text-to-video and image-to-video, ahead of previously leading models. It is one of the highest-rated video models on that leaderboard to date.

Open-source direction

Alibaba announced plans to release Happy Horse 1.0 under the permissive Apache-2.0 license, framing it as an open-source, state-of-the-art option. This positions it differently from closed commercial video models that are API-only.

Specs

Technical specifications

Model

NameHappy Horse 1.0 (HappyHorse-1.0)
DeveloperAlibaba, ATH-AI Innovation team
Parameters~15 billion
ArchitectureSingle-stream self-attention Transformer
License (announced)Apache-2.0

Generation

InputsText prompt, image
Output resolution1080p HD (native)
AudioNative synchronized audio (dialogue, Foley, ambience)
Multi-shotSupported within a single clip
Lip-sync languagesEnglish, Mandarin, Cantonese, Japanese, Korean, German, French

Recognition

LeaderboardArtificial Analysis Video Arena
Text-to-video rank#1 (blind human preference)
Image-to-video rank#1 (blind human preference)
DebutApril 2026 (anonymous, later attributed to Alibaba)
In practice

Use cases

Talking-character and dialogue clips

Because lip-sync and speech are generated with the video, the model suits avatars, explainers, and character scenes where mouth movement must match spoken lines — across seven languages without a separate voiceover pass.

Ready-to-watch social content

Native audio means short-form clips arrive with ambience and sound effects already attached, cutting the editing loop for creators publishing quickly to social platforms.

Localized video at scale

Multilingual lip-sync makes it practical to produce the same scene in several languages with matching mouth movement, useful for global marketing and education content.

Animating stills and concept art

Image-to-video lets designers and studios bring a single frame, product shot, or character illustration to life with motion and a synchronized soundtrack.

Prototyping and previsualization

Multi-shot 1080p generation with sound helps teams sketch out sequences, mood, and pacing early — a fuller preview than silent draft renders.

Generational leap

Happy Horse 1.0 vs. Seedance 2.0

Happy Horse 1.0 overtook ByteDance's Seedance 2.0 at the top of the Artificial Analysis Video Arena. Here's how the two leading models line up on the headline capabilities.

FeatureSeedance 2.0Happy Horse 1.0NEW
Video Arena rank (T2V, at HH debut)Previous leader#1
Native synchronized audioNot the core focusYes, joint audio-video in one pass
Multilingual lip-syncLimited7 languages
Announced open-source licenseClosed / commercialApache-2.0 (announced)
DeveloperByteDanceAlibaba
Honest look

Current limitations

Open-source weights still rolling out

Alibaba announced an Apache-2.0 release, but at the time of the sources reviewed the full model weights had not yet been publicly downloadable — the repository and model pages existed without complete artifacts. Availability may lag the announcement.

Benchmark figures vary by snapshot

Reported Elo scores and exact rankings differ between leaderboard captures and categories. Treat specific point values as approximate; the consistent takeaway is a top-tier, #1-class placement rather than one fixed number.

Clip length and per-shot constraints

Like most current video models, generations are short clips rather than long-form video, and complex multi-shot prompts can still produce continuity or artifact issues. Plan for iteration on demanding scenes.

Third-party access varies

Much of the early availability came through API partners and aggregators rather than a single official endpoint, so pricing, quotas, and feature parity can differ by provider until first-party access matures.

FAQ

Frequently asked questions

Who made Happy Horse 1.0?
It was developed by Alibaba. The model first appeared anonymously atop the Artificial Analysis Video Arena in April 2026, and Alibaba then confirmed it was built by an internal ATH-AI Innovation team led by Zhang Di, a former Kuaishou VP and technical architect of Kling AI.
What makes Happy Horse 1.0 different from other video models?
Its standout feature is native joint audio-video generation: a single ~15B-parameter Transformer produces the video and its synchronized soundtrack — dialogue, sound effects, and ambience — in one pass, instead of generating a silent clip and adding audio afterward.
What resolution and audio does it output?
It generates natively in 1080p HD with synchronized audio, including lip-synced dialogue. It supports multi-shot sequences and lip-sync across seven languages: English, Mandarin, Cantonese, Japanese, Korean, German, and French.
Is Happy Horse 1.0 open source?
Alibaba announced it would be released under the permissive Apache-2.0 license. Note that a license announcement is not the same as an immediate download — at the time of reporting the full weights had not yet been publicly released, so open availability may lag the announcement.
How good is it compared to other models?
On the Artificial Analysis Video Arena, which ranks models by blind human preference votes, Happy Horse 1.0 reached #1 in both text-to-video and image-to-video, ahead of previously leading models such as Seedance 2.0. Exact Elo figures vary by leaderboard snapshot.
Does it support both text and image inputs?
Yes. Happy Horse 1.0 supports text-to-video, generating a scene from a written prompt, and image-to-video, animating an existing still image or character design — both with the same native-audio pipeline.
Can I use Happy Horse 1.0 on SharkFoto?
Happy Horse 1.0 support on SharkFoto is coming soon. In the meantime you can use SharkFoto's currently available AI creative tools, and check back here for availability updates.

Happy Horse 1.0 is Coming to SharkFoto

We're adding Happy Horse 1.0 as platform access opens — explore SharkFoto's available AI tools today and check back for availability.

Try Happy Horse 1.0 now