MiniMax H3 — the model most people call Hailuo 3.0 — generates 4 to 15 second clips at 2K with native stereo audio produced in the same pass. Text, images, video and audio all flow in as one unified reference.
Dialogue, ambience and music are generated in the same pass as the picture, so lip movement, footsteps and score land in sync without a separate audio step.
Feed it up to 9 images, 3 videos and 3 audio clips in one request. Borrow camera motion from a video, a character from a photo and a voice from an audio clip, all directed in plain language.
Outputs 2K clips from 4 to 15 seconds in whole-second steps — long enough for a full beat of action, sharp enough to publish.
Start from a 7,000-character prompt, or drive motion from a first and last frame for tighter control over how a shot begins and ends.
MiniMax H3 — widely referred to by the community nickname Hailuo 3.0 — is MiniMax's omni-modal video model, released July 31, 2026 as the successor to the Hailuo 02 line. Its defining move is treating text, images, video and audio as a single unified context rather than separate pipelines. In one request you can describe a scene, hand it reference photos of a character, a clip whose camera move you want to copy and an audio track whose voice should carry the dialogue, then let the model resolve all of it into a coherent shot.
The headline capability is native sound. Where earlier Hailuo models produced silent footage that needed audio added afterward, H3 generates dialogue, ambience and music in the same pass as the picture, at 2K and durations from 4 to 15 seconds. MiniMax exposes three entry modes through its API — text-to-video, first/last-frame image-to-video, and omni-reference — and priced 2K generation at roughly $0.13 per second at launch. Open weights were promised shortly after release; treat availability of downloadable weights as forthcoming rather than confirmed.
H3 generates dialogue, ambient sound and music alongside the image in one generation, not as a post step. Because sound and picture come from the same pass, timing cues like speech, impacts and motion tend to stay aligned. This is the clearest break from Hailuo 02, which output silent video only.
A single request can carry up to 9 reference images, 3 reference videos and 3 reference audio clips — 12 conditioning inputs combined. You can lift camera movement from a video, hold a character's likeness from photos and match a voice from audio, all steered with natural-language instructions rather than separate control models.
Clips render at 2K in whole-second durations from 4 up to 15 seconds. Third-party reports cite roughly 24fps; MiniMax's own docs do not explicitly state the frame rate. The longer ceiling gives room for a complete action beat instead of a fragment.
Prompts accept up to 7,000 characters, enough to script shot composition, motion, lighting, dialogue and audio direction in detail. For text-to-video you explicitly choose an aspect ratio from a fixed set including 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16.
Image-to-video mode drives a shot from a starting frame, optionally with an ending frame, so you can pin how a clip opens and resolves. In image-to-video the output adapts to the source image's aspect ratio instead of requiring a manual ratio choice.
Beyond generation, H3 supports instruction-driven editing and motion transfer, letting you restyle or re-direct footage and carry a movement from a reference clip onto new subjects. Early independent benchmarks placed H3 strongly on video-editing tasks relative to peers.
At launch the model is reachable through MiniMax's API across its three entry modes. MiniMax stated open weights would follow within days of release; as of the launch coverage those weights had not yet appeared, so plan around the hosted API for now.
| Resolution | 2K |
|---|---|
| Duration | 4-15 seconds, integer only |
| Frame rate | ~24fps (third-party reported; not stated by MiniMax) |
| Audio | Native stereo, generated in-pass |
| Aspect ratios (text-to-video) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Entry modes | Text-to-video, first/last-frame image-to-video, omni-reference |
|---|---|
| Text prompt limit | Up to 7,000 characters |
| Reference images | Up to 9 per request |
| Reference videos | Up to 3 per request |
| Reference audio | Up to 3 per request |
| Image-to-video ratio | Adapts to source image |
| Developer | MiniMax |
|---|---|
| Official name / API id | MiniMax H3 / MiniMax-H3 |
| Common nickname | Hailuo 3.0 |
| Released | July 31, 2026 |
| Launch price | ~$0.13 / second at 2K |
| Open weights | Announced as forthcoming; not confirmed shipped at launch |
Because dialogue and lip motion come from the same pass, H3 suits short character scenes where a figure needs to speak on cue — explainers, ads or social skits — without stitching a voice track on afterward.
Hold a product or mascot steady with up to 9 reference images while borrowing a proven camera move from a reference video, so a series of clips shares a look and motion language across a campaign.
Feed a reference audio clip and let the picture respond to its mood and rhythm, useful for mood pieces, lyric-style visuals or atmospheric establishing shots where sound leads the edit.
Use first and last frame image-to-video to lock how a shot opens and closes — handy for transitions, logo reveals and match cuts where the exact start and end states matter.
Lean on instruction-based editing and motion transfer to re-direct or restyle existing footage, or map a movement from one clip onto a new subject, without rebuilding a scene from scratch.
H3 is the successor to MiniMax's Hailuo 02. The jump is less about raw resolution and more about becoming omni-modal — folding audio and multi-source references into a single generation.
| Feature | Hailuo 02 | MiniMax H3 (Hailuo 3.0)NEW |
|---|---|---|
| Max resolution | Up to 1080p | 2K |
| Max duration | 6-10 seconds | 4-15 seconds |
| Native audio | No — silent output | Yes — stereo, same pass |
| Reference inputs | Image-to-video / text-to-video | Omni-reference: up to 9 images, 3 videos, 3 audio |
| Editing & motion transfer | Not a core feature | Instruction-based editing and motion transfer |
| Released | 2025 | July 31, 2026 |
MiniMax said H3's weights would arrive within days of the July 31 launch, but as of the launch coverage no repository had appeared. Until they ship, the hosted API is the only reliable path, so self-hosting plans should stay provisional.
Independent reports consistently cite ~24fps, but MiniMax's own documentation does not explicitly confirm a frame rate. Treat 24fps as likely rather than guaranteed if exact motion cadence matters to your pipeline.
Maximum duration is 15 seconds in whole-second steps. H3 is built for short, high-density clips; longer narratives still need to be assembled from multiple generations.
Early third-party benchmarks put H3 ahead on video editing but trailing some rivals on pure text-to-video and image-to-video quality. It is a strong all-rounder with native audio, not a category leader on every single metric.
We're adding MiniMax H3 (Hailuo 3.0) as platform access opens — explore SharkFoto's available AI tools today and check back for availability.
Try MiniMax H3 (Hailuo 3.0) now