MiniMax · Omni-modal video with sound baked in.

MiniMax H3 (Hailuo 3.0)

One model. Picture and sound, together.

MiniMax H3 — the model most people call Hailuo 3.0 — generates 4 to 15 second clips at 2K with native stereo audio produced in the same pass. Text, images, video and audio all flow in as one unified reference.

2K
Max resolution
15s
Max duration
Stereo
Native audio
12
Reference inputs
What's new

One model. Picture and sound, together.

01

Native audio, not a bolt-on

Dialogue, ambience and music are generated in the same pass as the picture, so lip movement, footsteps and score land in sync without a separate audio step.

02

Omni-reference input

Feed it up to 9 images, 3 videos and 3 audio clips in one request. Borrow camera motion from a video, a character from a photo and a voice from an audio clip, all directed in plain language.

03

2K, up to 15 seconds

Outputs 2K clips from 4 to 15 seconds in whole-second steps — long enough for a full beat of action, sharp enough to publish.

04

Text or image to video

Start from a 7,000-character prompt, or drive motion from a first and last frame for tighter control over how a shot begins and ends.

Overview

Model at a glance

MiniMax H3 — widely referred to by the community nickname Hailuo 3.0 — is MiniMax's omni-modal video model, released July 31, 2026 as the successor to the Hailuo 02 line. Its defining move is treating text, images, video and audio as a single unified context rather than separate pipelines. In one request you can describe a scene, hand it reference photos of a character, a clip whose camera move you want to copy and an audio track whose voice should carry the dialogue, then let the model resolve all of it into a coherent shot.

The headline capability is native sound. Where earlier Hailuo models produced silent footage that needed audio added afterward, H3 generates dialogue, ambience and music in the same pass as the picture, at 2K and durations from 4 to 15 seconds. MiniMax exposes three entry modes through its API — text-to-video, first/last-frame image-to-video, and omni-reference — and priced 2K generation at roughly $0.13 per second at launch. Open weights were promised shortly after release; treat availability of downloadable weights as forthcoming rather than confirmed.

Official name
MiniMax H3 (API id MiniMax-H3)
Common nickname
Hailuo 3.0
Vendor
MiniMax
Modality
Omni-modal to video
Resolution
2K
Duration
4-15s, integer steps
Native audio
Stereo, generated in-pass
Launch price
~$0.13 / second at 2K
Capabilities

Key features

Native stereo audio in a single pass

H3 generates dialogue, ambient sound and music alongside the image in one generation, not as a post step. Because sound and picture come from the same pass, timing cues like speech, impacts and motion tend to stay aligned. This is the clearest break from Hailuo 02, which output silent video only.

Omni-reference across four modalities

A single request can carry up to 9 reference images, 3 reference videos and 3 reference audio clips — 12 conditioning inputs combined. You can lift camera movement from a video, hold a character's likeness from photos and match a voice from audio, all steered with natural-language instructions rather than separate control models.

2K output, 4 to 15 seconds

Clips render at 2K in whole-second durations from 4 up to 15 seconds. Third-party reports cite roughly 24fps; MiniMax's own docs do not explicitly state the frame rate. The longer ceiling gives room for a complete action beat instead of a fragment.

Text-to-video with long prompts

Prompts accept up to 7,000 characters, enough to script shot composition, motion, lighting, dialogue and audio direction in detail. For text-to-video you explicitly choose an aspect ratio from a fixed set including 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16.

First and last frame image-to-video

Image-to-video mode drives a shot from a starting frame, optionally with an ending frame, so you can pin how a clip opens and resolves. In image-to-video the output adapts to the source image's aspect ratio instead of requiring a manual ratio choice.

Instruction-based editing and motion transfer

Beyond generation, H3 supports instruction-driven editing and motion transfer, letting you restyle or re-direct footage and carry a movement from a reference clip onto new subjects. Early independent benchmarks placed H3 strongly on video-editing tasks relative to peers.

API-first, with weights promised

At launch the model is reachable through MiniMax's API across its three entry modes. MiniMax stated open weights would follow within days of release; as of the launch coverage those weights had not yet appeared, so plan around the hosted API for now.

Specs

Technical specifications

Output

Resolution2K
Duration4-15 seconds, integer only
Frame rate~24fps (third-party reported; not stated by MiniMax)
AudioNative stereo, generated in-pass
Aspect ratios (text-to-video)21:9, 16:9, 4:3, 1:1, 3:4, 9:16

Inputs & modes

Entry modesText-to-video, first/last-frame image-to-video, omni-reference
Text prompt limitUp to 7,000 characters
Reference imagesUp to 9 per request
Reference videosUp to 3 per request
Reference audioUp to 3 per request
Image-to-video ratioAdapts to source image

Model & access

DeveloperMiniMax
Official name / API idMiniMax H3 / MiniMax-H3
Common nicknameHailuo 3.0
ReleasedJuly 31, 2026
Launch price~$0.13 / second at 2K
Open weightsAnnounced as forthcoming; not confirmed shipped at launch
In practice

Use cases

Talking-character shorts

Because dialogue and lip motion come from the same pass, H3 suits short character scenes where a figure needs to speak on cue — explainers, ads or social skits — without stitching a voice track on afterward.

Reference-consistent brand clips

Hold a product or mascot steady with up to 9 reference images while borrowing a proven camera move from a reference video, so a series of clips shares a look and motion language across a campaign.

Music and ambience-driven scenes

Feed a reference audio clip and let the picture respond to its mood and rhythm, useful for mood pieces, lyric-style visuals or atmospheric establishing shots where sound leads the edit.

Controlled shot bookends

Use first and last frame image-to-video to lock how a shot opens and closes — handy for transitions, logo reveals and match cuts where the exact start and end states matter.

Quick edits and restyles

Lean on instruction-based editing and motion transfer to re-direct or restyle existing footage, or map a movement from one clip onto a new subject, without rebuilding a scene from scratch.

Generational leap

Hailuo 02 vs MiniMax H3

H3 is the successor to MiniMax's Hailuo 02. The jump is less about raw resolution and more about becoming omni-modal — folding audio and multi-source references into a single generation.

FeatureHailuo 02MiniMax H3 (Hailuo 3.0)NEW
Max resolutionUp to 1080p2K
Max duration6-10 seconds4-15 seconds
Native audioNo — silent outputYes — stereo, same pass
Reference inputsImage-to-video / text-to-videoOmni-reference: up to 9 images, 3 videos, 3 audio
Editing & motion transferNot a core featureInstruction-based editing and motion transfer
Released2025July 31, 2026
Honest look

Current limitations

Open weights not yet confirmed

MiniMax said H3's weights would arrive within days of the July 31 launch, but as of the launch coverage no repository had appeared. Until they ship, the hosted API is the only reliable path, so self-hosting plans should stay provisional.

Frame rate not officially stated

Independent reports consistently cite ~24fps, but MiniMax's own documentation does not explicitly confirm a frame rate. Treat 24fps as likely rather than guaranteed if exact motion cadence matters to your pipeline.

Short-form ceiling

Maximum duration is 15 seconds in whole-second steps. H3 is built for short, high-density clips; longer narratives still need to be assembled from multiple generations.

Not the strongest on every axis

Early third-party benchmarks put H3 ahead on video editing but trailing some rivals on pure text-to-video and image-to-video quality. It is a strong all-rounder with native audio, not a category leader on every single metric.

FAQ

Frequently asked questions

Is it called Hailuo H3 or MiniMax H3?
MiniMax's official name is MiniMax H3, with the API model id MiniMax-H3. Many people call it Hailuo 3.0 because it succeeds the Hailuo video line, and you'll also see the hybrid Hailuo H3. They all refer to the same July 31, 2026 model.
Does MiniMax H3 generate audio?
Yes. Unlike Hailuo 02, which output silent video, H3 generates native stereo audio — dialogue, ambience and music — in the same pass as the picture, so sound and motion are aligned by default.
What resolution and length can it produce?
H3 outputs 2K video in durations from 4 to 15 seconds, in whole-second steps. Third-party sources report roughly 24fps, though MiniMax does not explicitly state the frame rate.
What can I reference in a single request?
Its omni-reference system accepts up to 9 images, 3 videos and 3 audio clips — 12 conditioning inputs combined — letting you borrow a character, a camera move and a voice at once, directed in natural language.
Does it support image-to-video?
Yes. H3 offers first/last-frame image-to-video, where the output adapts to the source image's aspect ratio, as well as text-to-video from prompts up to 7,000 characters.
Are the model weights open?
MiniMax announced that H3's weights would be released within days of launch, but they had not yet appeared as of the launch coverage. For now the hosted API is the dependable way to access the model.
Can I use MiniMax H3 on SharkFoto?
MiniMax H3 (Hailuo 3.0) support on SharkFoto is coming soon. In the meantime you can use SharkFoto's currently available AI creative tools, and check back here for availability updates.

MiniMax H3 (Hailuo 3.0) is Coming to SharkFoto

We're adding MiniMax H3 (Hailuo 3.0) as platform access opens — explore SharkFoto's available AI tools today and check back for availability.

Try MiniMax H3 (Hailuo 3.0) now