Unified multimodal context
One request can carry text alongside up to nine images, three videos of up to 15 seconds and three audio clips of up to 15 seconds, so the model reads the whole brief rather than a prompt string.
MiniMax (Open Weights)
Open-weights unified multimodal video model. Up to 15 seconds at 2K with native stereo audio, driven by a unified context that accepts text, images, video and audio in the same request. The weights are open, so it can also be run on your own hardware.
The base MiniMax H3 endpoints are not in the Studio catalog yet — these buttons open our text-to-video and image-to-video workflows, and the speed-optimised MiniMax H3 Max is available in the Studio today at /models/minimax-h3-max.
Features
One request can carry text alongside up to nine images, three videos of up to 15 seconds and three audio clips of up to 15 seconds, so the model reads the whole brief rather than a prompt string.
Swap a product, rewrite signage, relight a scene or add and remove an element while the rest of the frame stays where it was.
Legible type and animated interface work hold up at 2K, and prompts run to 7000 characters when a shot needs that much direction.
Audio is generated in stereo in the same pass as the picture, and a supplied voice can be transferred or cloned onto the speaking character.
The weights are published alongside the hosted release, so the same model can be run on your own infrastructure.
Examples
Vintage binocular brand film, Images 1 to 4 as sequential keyframes, Wes Anderson 35mm look, red typographic accents
Try this promptWuxia character film, Image 2 locked reference (hanfu and silver crown), 4K xianxia production value
Try this promptLive-action to voxel: preserve buildings, transform only trees and cars to Minecraft blocks
Try this promptNeon laundromat encounter, live-action plus hand-drawn animation, phone-camera feel
Try this promptCapybara recreation: replace three suited men with photoreal capybaras, exact movement path
Try this promptGreen-screen fairytale composite: remove green screen, match background to characters' actions
Try this promptThese clips and prompts are published by MiniMax as reference for what H3 can do. Run one in the Studio to see your own render.
Workflows
MiniMax has not published a public rate for H3, so there is no per-second figure to quote. Credits come from one shared balance across image, video and speech, and the exact estimate is shown in the Studio before you generate.
Built for
FAQ
Limits, audio, licensing and price — the things worth knowing before you spend a credit.
Open the Studio to start in text-to-video, or run the speed-optimised MiniMax H3 Max at /models/minimax-h3-max today.