MiniMax H3

MiniMax-H3

A 33B-parameter omni-modal model that generates video and a fully synchronized soundtrack — up to 15 seconds, 2K resolution, 24 FPS, stereo audio.

Watch it work

Official example generations from the model repository.

TEXT → VIDEO

Text to video + audio

A starship fleet jumping to hyperspace, with music, ambience and foley generated alongside the picture.

IMAGE → VIDEO

First/last-frame mode

Animate from one keyframe (or two, to bookend the shot) while keeping the subject consistent.

REFERENCE → VIDEO

Omni-reference mode

Guide generation with up to 9 reference images, 3 video clips and audio — in any mix.

2K REGENERATION

2K upscale pass

The same scene regenerated at 2K via H3-Regenerate, recovering fine detail and small text.

The three modes

VariantInputWhat you get
T2VAText promptVideo + soundtrack from nothing but words
FL2VAText + first frame, last frame, or bothAnimate between keyframes with consistent subjects
Ref2VAText + images / videos / audio (≤12 files)Reference-guided generation: match a look, a voice, a scene

Prompt guidance

H3 follows structured, cinematic prompts well — describe shots, camera moves, sound and music explicitly:

"Cinematic, medium wide shot, pushing in slowly. A red fox trotting through a snowy pine forest at dawn, snow crunching underfoot. Low resonant forest ambience, soft wind through the pines, swelling warm orchestral score."
"A busy night market, neon signs reflecting in puddles, sizzling street food. Shallow depth of field, handheld feel. Bustling crowd chatter, sizzling oil, distant city hum."

Try it yourself

Generate your own video with the interactive demo on ZeroGPU (visitors use their own free GPU quota):

Or call the official API at platform.minimax.io · MIT-licensed usage per the MiniMax H3 Community License.