1 link tagged with all of: text-to-video + ai-model + video-ai
Click any tag below to further narrow down your results
Links
MiniMax H3 is a general-purpose AI that takes text, images, video, and audio as input to generate up to 15-second, 2K videos with native stereo sound. It matches or beats mainstream models on price-performance, excels at instruction following and brand rendering, and will release its weights soon under open-source terms.
- MiniMax H3 unifies text, image, video, and audio into one model that generates 15-second 2K videos with native stereo sound, costing under one-third the price per second at 2K versus mainstream models.
- It can follow complex multimodal instructions, merging camera movement, character appearance, and audio from separate source inputs into one coherent output.
- Four new technologies—Contextual Omni Representation, H3-VAE, Omni Transformer, and In-Context Regeneration—drive its compression, throughput (+30%), and cross-modal coherence.
- MiniMax plans to open-source the model weights, designed for broad hardware compatibility, aiming to seed a wider ecosystem.