|

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

MiniMax releases MiniMax H3, a general-purpose multimodal technology mannequin. MiniMax H3 shouldn’t be a text-to-video mannequin with add-ons. MiniMax describes it as a general-purpose multimodal technology mannequin that reads textual content, pictures, video, and audio as one unified context and returns video with native stereo sound. The mains specs embody: 2K output, 4–15 seconds, integer durations solely.

Previous video stacks break up into text-to-video, image-to-video, first-and-last-frame, topic reference, movement reference, and video modifying, every typically a separate professional mannequin. MiniMax H3 folds these into one pretraining paradigm the place reference and modifying relationships are expressed in pure language. MiniMax’s instance immediate makes the purpose: reference the digicam motion from Video 1, have the character in Image 2 sing, match the vocals to Audio 3.

Is it deployable?

Today: sure, by way of the API and no, by yourself {hardware}. MiniMax launched H3 on July 31, 2026 with the mannequin dwell within the platform API underneath the mannequin ID MiniMax-H3 and within the shopper Hailuo AI app.

Industries: MiniMax positions MiniMax H3 for promoting, branding, e-commerce, product design, UI/UX, and gaming together with movie pre-visualization and retail catalog media.

Applications: Ad variant technology, product and itemizing movies, animated posters, movie title sequences, web site hero loops, character-consistent recreation cinematics, and video-to-video movement switch.

The API floor

The video generation guide paperwork three entry modes: text-to-video, first/last-frame image-to-video, and reference technology. Behind one endpoint and an asynchronous three-step movement: create a activity, ballot task_id, obtain content material.url.

Input limits value designing round:

  • Reference pictures: as much as 9. Reference movies: as much as 3 clips, 2–15 s every, ≤15 s whole. Reference audio: as much as 3 clips, and audio can’t be despatched with out an accompanying picture or video.
  • Mixed enter caps at 12 recordsdata whole. Prompt size ≤7,000 characters; request physique ≤64 MB, with URL enter advisable for giant property.
  • File sizes: video ≤50 MB, picture ≤30 MB, audio ≤15 MB, per asset.
  • Formats: H.264/H.265 video, JPG/PNG/WEBP/HEIC/HEIF pictures, WAV/MP3 audio.

Four technical items doing the work

Contextual Omni Representation: MiniMax rebuilt captioning so it describes the relationship between context and goal video, not simply the goal. Most supply materials requires roughly 100K tokens of inference, distilled to about 4K tokens on common. Language is the bridge that turns a hard and fast activity set into an open, descriptive one.

H3-VAE: A full tokenizer overhaul. Its excessive compression ratio delivers a said 4× achieve in efficient sequence size, reducing coaching and inference value and it’s the enabling expertise for native 2K.

H3-Omni Transformer: MiniMax explicitly put aside the Hailuo-02 structure right here. Multimodal context tripled sequence-length variance, so the coaching structure separates understanding and technology workloads and tunes {hardware} utilization for every. Reported end result: end-to-end coaching throughput up almost 30%.

In-Context Regeneration: Instead of a bolt-on super-resolution module, the bottom mannequin regenerates its personal low-resolution output in-context, re-reading the unique multimodal context. That is what recovers small textual content and wonderful element that typical upscalers guess at — immediately related to model and product rendering.


Check out the Technical detailsAlso, be at liberty to comply with us on Twitter and don’t overlook to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to accomplice with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so forth.? Connect with us

The submit MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio appeared first on MarkTechPost.

Similar Posts