Click any tag below to further narrow down your results
Links
This repo introduces LongCat-Video, a 13.6B-parameter model that handles text-to-video, image-to-video, and video continuation within a single framework. It uses block sparse attention and a coarse-to-fine strategy to produce minutes-long 720p/30fps videos without quality drift. The project also includes an audio-driven avatar extension with Whisper-based lip sync and distillation-accelerated inference.
- A single 13.6B-parameter model handles text-to-video, image-to-video, and video continuation, producing minutes-long 720p/30fps clips without quality drift.
- Avatar 1.5 (May 2026) switched from wav2vec2 to Whisper-Large-v3 for better lip sync, improved long-video stability, and added stylized domains like anime and animal content.
- Step-distilled inference cuts video generation down to just eight steps, with INT8 quantization also supported.
MiniMax H3 is a general-purpose AI that takes text, images, video, and audio as input to generate up to 15-second, 2K videos with native stereo sound. It matches or beats mainstream models on price-performance, excels at instruction following and brand rendering, and will release its weights soon under open-source terms.
- MiniMax H3 unifies text, image, video, and audio into one model that generates 15-second 2K videos with native stereo sound, costing under one-third the price per second at 2K versus mainstream models.
- It can follow complex multimodal instructions, merging camera movement, character appearance, and audio from separate source inputs into one coherent output.
- Four new technologies—Contextual Omni Representation, H3-VAE, Omni Transformer, and In-Context Regeneration—drive its compression, throughput (+30%), and cross-modal coherence.
- MiniMax plans to open-source the model weights, designed for broad hardware compatibility, aiming to seed a wider ecosystem.
This article covers a service that turns plain text into HD cartoon videos in just a few minutes. It uses AI to generate characters, compose scenes, animate movements, and export the final clip—all without any animation skills.
- Generates HD cartoon videos from text descriptions in about three minutes
- Four-step workflow: write description, AI creates characters, AI animates scenes, export video
- Maintains character consistency across scenes and supports multiple aspect ratios for different platforms
- Requires no drawing tools or animation experience—just a text prompt
+ ai-animation
text-to-video
+ video-generation
+ cartoon
+ character-generation
+ text
+ to
+ video