# MiniMax H3 — research findings Compiled 2026-08-06. Sources linked inline; benchmark numbers we measured ourselves live in `bench/results.csv` and `docs/BENCHMARKS.md`. ## The model [MiniMax H3](https://www.minimax.io/blog/minimax-h3) is an open-weights omni-modal video model: one 33B dense "Omni-Transformer" DiT (50 layers, hidden 5376) jointly denoises video + **native stereo audio** latents in a single forward pass, conditioned on hidden states from a **Qwen3-VL-32B** text encoder (layer 50). CFG-distilled — no negative prompt / CFG pass. Output is 24 fps fixed, 4–15 s, native canvas ~1 MP (2K comes from a separate regenerate/upscale module in the hosted API). Video VAE: 16x spatial / 4x temporal. Audio VAE: stereo 32 kHz, 40 latent fps. Two task checkpoints, each shipped in several precisions: - **fl2va** — text-to-video and first/last-frame conditioning (covers t2v + i2v) - **ref2va** — reference-to-video: up to 9 ref images, 3 ref videos (with soundtracks), 3 standalone audio clips, woven in via `` / `