MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities - MiniMax Research
MiniMax launched H3, a general-purpose multimodal generation model that generates video with native stereo sound, up to 15 seconds at 2K resolution, with open weights planned.
- H3 generates video with native stereo sound, up to 15 seconds at 2K resolution.
- At 2K, H3's per-second price is less than a third of mainstream models; at 768p, less than half the price of mainstream models' 720p.
- H3 supports multimodal context understanding, instruction following, and V2V motion transfer for commercial content creation.
- Model weights will be open-sourced in the coming days, subject to applicable laws and regulations.
- H3 uses H3-VAE, H3-Omni Transformer, and In-Context Regeneration technologies.
Content creators and developers can now use an open, cost-effective multimodal model for tasks like advertising, branding, and e-commerce, reducing reliance on closed-source video generation models.
Sources
Read this as text
Back to the AI news