paper-with-me

Papers

LongCat-Video Technical Report

2025-10-25 · Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang, Hongyu Li, Shijun Liang, Liya Ma, Siyu Ren, Xiaoming Wei, Rixu Xie, Tong Zhang arxiv

Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameters, delivering strong performance across multiple video generation tasks. It particularly excels in efficient and high-quality long video generation, representing our first step toward world models. Key features include: Unified architecture for multiple tasks: Built on the Diffusion Transformer (DiT) framework, LongCat-Video supports Text-to-Video, Image-to-Video, and Video-Continuation tasks with a single model; Long video generation: Pretraining on Video-Continuation tasks enables LongCat-Video to maintain high quality and temporal coherence in the generation of minutes-long videos; Efficient inference: LongCat-Video generates 720p, 30fps videos within minutes by employing a coarse-to-fine generation strategy along both the temporal and spatial axes. Block Sparse Attention further enhances efficiency, particularly at high resolutions; Strong performance with multi-reward RLHF: Multi-reward RLHF training enables LongCat-Video to achieve performance on par with the latest closed-source and leading open-source models. Code and model weights are publicly available to accelerate progress in the field.

📄 PDF Abstract BibTeX arXiv:2510.22200

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

LongCat-Flash-Omni Technical Report

2025-10-31 · Meituan LongCat Team, Bairui Wang, Bayan, Bin Xiao 외 arxiv

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strate…

LongCat-Video-Avatar 1.5 Technical Report

2026-05-26 · Meituan LongCat Team, Xunliang Cai, Meng Cheng, Feng Gao 외 arxiv

Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework prioritizing systematic engineering and …

Video Generation

LongCat-Flash Technical Report

2025-09-01 · Meituan LongCat Team, Bayan, Bei Li, Bingye Lei 외 arxiv

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, Lo…

Computational Efficiency

Introducing LongCat-Flash-Thinking: A Technical Report

2025-09-23 · Meituan LongCat Team, Anchun Gui, Bei Li, Bingyang Tao 외 arxiv

We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginn…

Reinforcement Learning

LongCat-Image Technical Report

2025-12-08 · Meituan LongCat Team, Hanghang Ma, Haoxian Tan, Jiale Huang 외 arxiv

We introduce LongCat-Image, a pioneering open-source and bilingual (Chinese-English) foundation model for image generation, designed to address core challenges in multilingual text rendering, photorealism, deployment eff…

Image GenerationImage Editing