MiniMax Releases H3 Multimodal Video Model with Upcoming Open Weights
MiniMax has launched MiniMax-H3, a multimodal model capable of generating up to 15-second videos at 2K resolution with native stereo sound. The company also announced plans to release the model's weights to the open-source community in the coming days. This release challenges the dominance of closed-source video generation models by providing a high-performance, cost-effective open alternative. It enables developers to customize the model and run it on a wider range of AI hardware. The model is powered by technologies like H3-VAE and H3-Omni Transformer, excelling in instruction following, brand rendering, and video-to-video motion transfer. It is currently available in the Video Arena for text-to-video and image-to-video testing.
## BACKGROUND
Video-to-video motion transfer is an AI technique that extracts motion from a reference video and applies it to a target character or image. Native multimodality, as previously explored in models like MiniMax M3, allows a single neural network to process and generate multiple data types like text, audio, and video simultaneously.