MiniMax to Open-Source H3 Multimodal Video Model Supporting 15s 2K Resolution
MiniMax has announced that its H3 general multimodal video model will be officially open-sourced on the ModelScope platform on August 3. The model is capable of generating up to 15-second videos at 2K resolution with native dual-channel audio. By offering 2K video generation at less than one-third the cost of mainstream models, MiniMax H3 significantly lowers the financial barrier for high-quality AI video production. This cost efficiency, combined with native audio integration, could accelerate the adoption of generative AI in advertising, gaming, and filmmaking. The model leverages advanced technologies such as Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-context Regeneration to achieve high performance. It supports precise multi-dimensional editing and control over characters, objects, scenes, and sound, while maintaining strict instruction-following capabilities.
## BACKGROUND
MiniMax is a prominent Chinese artificial intelligence startup specializing in large language and multimodal models. ModelScope is an open-source model community and platform initiated by Alibaba to host and share advanced machine learning models.