MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio
MiniMax releases MiniMax H3, a general-purpose multimodal technology mannequin. MiniMax H3 shouldn’t be a text-to-video mannequin with add-ons. MiniMax describes it as a general-purpose multimodal technology mannequin that reads textual content, pictures, video, and audio as one unified context and returns video with native stereo sound. The mains specs embody: 2K output, 4–15 seconds, integer durations solely.
Previous video stacks break up into text-to-video, image-to-video, first-and-last-frame, topic reference, movement reference, and video modifying, every typically a separate professional mannequin. MiniMax H3 folds these into one pretraining paradigm the place reference and modifying relationships are expressed in pure language. MiniMax’s instance immediate makes the purpose: reference the digicam motion from Video 1, have the character in Image 2 sing, match the vocals to Audio 3.
Is it deployable?
Today: sure, by way of the API and no, by yourself {hardware}. MiniMax launched H3 on July 31, 2026 with the mannequin dwell within the platform API underneath the mannequin ID MiniMax-H3 and within the shopper Hailuo AI app.
Industries: MiniMax positions MiniMax H3 for promoting, branding, e-commerce, product design, UI/UX, and gaming together with movie pre-visualization and retail catalog media.
Applications: Ad variant technology, product and itemizing movies, animated posters, movie title sequences, web site hero loops, character-consistent recreation cinematics, and video-to-video movement switch.
The API floor
The video generation guide paperwork three entry modes: text-to-video, first/last-frame image-to-video, and reference technology. Behind one endpoint and an asynchronous three-step movement: create a activity, ballot task_id, obtain content material.url.
Input limits value designing round:
- Reference pictures: as much as 9. Reference movies: as much as 3 clips, 2–15 s every, ≤15 s whole. Reference audio: as much as 3 clips, and audio can’t be despatched with out an accompanying picture or video.
- Mixed enter caps at 12 recordsdata whole. Prompt size ≤7,000 characters; request physique ≤64 MB, with URL enter advisable for giant property.
- File sizes: video ≤50 MB, picture ≤30 MB, audio ≤15 MB, per asset.
- Formats: H.264/H.265 video, JPG/PNG/WEBP/HEIC/HEIF pictures, WAV/MP3 audio.
Four technical items doing the work
Contextual Omni Representation: MiniMax rebuilt captioning so it describes the relationship between context and goal video, not simply the goal. Most supply materials requires roughly 100K tokens of inference, distilled to about 4K tokens on common. Language is the bridge that turns a hard and fast activity set into an open, descriptive one.
H3-VAE: A full tokenizer overhaul. Its excessive compression ratio delivers a said 4× achieve in efficient sequence size, reducing coaching and inference value and it’s the enabling expertise for native 2K.
H3-Omni Transformer: MiniMax explicitly put aside the Hailuo-02 structure right here. Multimodal context tripled sequence-length variance, so the coaching structure separates understanding and technology workloads and tunes {hardware} utilization for every. Reported end result: end-to-end coaching throughput up almost 30%.
In-Context Regeneration: Instead of a bolt-on super-resolution module, the bottom mannequin regenerates its personal low-resolution output in-context, re-reading the unique multimodal context. That is what recovers small textual content and wonderful element that typical upscalers guess at — immediately related to model and product rendering.
