A 33B-parameter omni-modal model that generates video and a fully synchronized soundtrack — up to 15 seconds, 2K resolution, 24 FPS, stereo audio.
Official example generations from the model repository.
A starship fleet jumping to hyperspace, with music, ambience and foley generated alongside the picture.
Animate from one keyframe (or two, to bookend the shot) while keeping the subject consistent.
Guide generation with up to 9 reference images, 3 video clips and audio — in any mix.
The same scene regenerated at 2K via H3-Regenerate, recovering fine detail and small text.
| Variant | Input | What you get |
|---|---|---|
| T2VA | Text prompt | Video + soundtrack from nothing but words |
| FL2VA | Text + first frame, last frame, or both | Animate between keyframes with consistent subjects |
| Ref2VA | Text + images / videos / audio (≤12 files) | Reference-guided generation: match a look, a voice, a scene |
H3 follows structured, cinematic prompts well — describe shots, camera moves, sound and music explicitly:
Generate your own video with the interactive demo on ZeroGPU (visitors use their own free GPU quota):
Or call the official API at platform.minimax.io · MIT-licensed usage per the MiniMax H3 Community License.