Video generation with built-in stereo sound in a single model: MiniMax H3 reads text, images, video and audio as one context and returns 2K clips of 4 to 15 seconds with native audio.
The model replaces the usual collection of expert pipelines, such as separate text-to-video, image-to-video, frame-reference and editing models, with a single pre-training paradigm where reference and editing relationships are expressed in natural language. A prompt can borrow camera movement from one clip, place a character from an image and match vocals from an audio file.
H3 went live on July 31 through the platform API as MiniMax-H3 and in the Hailuo AI consumer app, aimed at advertising, branding, e-commerce, gaming and film pre-visualization.
Under the hood sit four components: captioning that describes relationships between context and target, compressing roughly 100K tokens of inference to about 4K; an H3-VAE tokenizer claiming 4x effective sequence length that enables native 2K; an H3-Omni Transformer separating understanding and generation workloads for close to 30% higher training throughput; and in-context regeneration that re-reads the original context rather than applying a super-resolution module.
Input limits cap reference images at 9, reference videos at 3 clips of 2-15 seconds, and audio at 3 clips, with mixed input limited to 12 files and prompts to 7,000 characters.