MiniMax-H3 Ties for Third in Text-to-Video Arena Benchmark
The newly released MiniMax-H3 model has achieved a score of 1,455 points in the Text-to-Video Arena benchmark, tying for third place overall and sitting just three points behind Muse Video. Additionally, it outperformed the next best open model by more than 283 points. This benchmark result positions MiniMax-H3 in the top tier of video generation models, demonstrating the rapid advancement of open, multimodal models in matching or exceeding proprietary alternatives. MiniMax-H3 is an open omni-modal model capable of processing text, image, video, and audio inputs to generate up to 2K resolution videos with native stereo audio lasting up to 15 seconds. Other models from the same developer, such as Hailuo-2.3 and Hailuo-02-pro, ranked 23rd and 29th respectively.
## BACKGROUND
The Text-to-Video Arena is a blind-test benchmark platform where users compare anonymous video clips generated by different AI models using the same prompt and vote on the better output. MiniMax is an AI company known for its generative models, and its H3 release represents a shift toward unified multimodal intelligence rather than specialized single-task generation.