More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
LongCat-Video packs a 13.6 billion-parameter transformer into one model that handles text-to-video, image-to-video and video-continuation tasks. Its backbone uses a coarse-to-fine strategy along both time and space axes, plus block-sparse attention, to crank out 720 p, 30 fps clips in minutes. Pretraining on video-continuation makes it resilient to color drift and quality loss, even over minutes-long outputs. Multi-reward Group Relative Policy Optimization (GRPO) fine-tunes it via RLHF, pushing its results on public and internal benchmarks up to par with leading open-source and commercial systems.
Since the initial October 2025 release, the project launched two avatar-focused versions. LongCat-Video-Avatar debuted in December 2025 for audio-driven character animation, covering Audio-Text-to-Video, Audio-Text-Image-to-Video and continuation modes. In May 2026, Avatar 1.5 swapped wav2vec2 for Whisper-Large-v3, boosted lip sync, improved long-video stability, and added stylized domains (anime, animal, complex scenes). Step-distilled inference cuts generation to eight steps. The GitHub repo includes conda setup commands, FlashAttention-2 installation, demo scripts for single- and multi-GPU runs, plus huggingface-cli commands to pull weights for all three models. User notes cover audio CFG settings, prompt detail, reference-image indexing, mask-frame ranges, resolution flags, dual-audio modes, distillation and INT8 quantization options.
Questions about this article
No questions yet.