MiniMax Launches H3 All-Modal Generative Model with 2K Video Support
MiniMax has officially launched its H3 all-modal generative model, capable of generating up to 15-second 2K videos with native stereo sound. The model utilizes a novel in-context regeneration approach for video upscaling and its weights will be open-sourced soon. By offering 2K video generation at less than a third of the price of mainstream models, H3 lowers the barrier to high-quality content creation. Furthermore, open-sourcing the model weights will challenge the dominance of closed-source video models and foster community-driven customization. Instead of traditional super-resolution modules, H3 uses its base model to regenerate low-resolution outputs in-context, recovering details that traditional methods cannot. It also excels at instruction following, brand rendering, and video-to-video (V2V) motion transfer.
## BACKGROUND
Traditional video upscaling relies on separate super-resolution algorithms that often struggle to accurately reconstruct fine details, sometimes leading to artifacts. Video-to-video (V2V) motion transfer is a technique that applies the motion and dynamics of a source video to a different character or image.