paper-with-me

Papers

LTX-2: Efficient Joint Audio-Visual Foundation Model

2026-01-06 · Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, Eitan Richardson, Guy Shiran, Itay Chachy, Jonathan Chetboun, Michael Finkelson, Michael Kupchick, Nir Zabari, Nitzan Guetta, Noa Kotler, Ofir Bibi, Ori Gordon, Poriya Panet, Roi Benita, Shahar Armon, Victor Kulikov, Yaron Inger, Yonatan Shiftan, Zeev Melumian, Zeev Farbman arxiv

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model capable of generating high-quality, temporally synchronized audiovisual content in a unified manner. LTX-2 consists of an asymmetric dual-stream transformer with a 14B-parameter video stream and a 5B-parameter audio stream, coupled through bidirectional audio-video cross-attention layers with temporal positional embeddings and cross-modality AdaLN for shared timestep conditioning. This architecture enables efficient training and inference of a unified audiovisual model while allocating more capacity for video generation than audio generation. We employ a multilingual text encoder for broader prompt understanding and introduce a modality-aware classifier-free guidance (modality-CFG) mechanism for improved audiovisual alignment and controllability. Beyond generating speech, LTX-2 produces rich, coherent audio tracks that follow the characters, environment, style, and emotion of each scene -- complete with natural background and foley elements. In our evaluations, the model achieves state-of-the-art audiovisual quality and prompt adherence among open-source systems, while delivering results comparable to proprietary models at a fraction of their computational cost and inference time. All model weights and code are publicly released.

📄 PDF Abstract BibTeX arXiv:2601.03233

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationVideo Generation

Similar Papers 제목 키워드 기반

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

2026-01-29 · Anthony Chen, Naomi Ken Korem, Gal Zeevi, Tavi Halperin 외 arxiv

Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for d…

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

2025-12-15 · Team Seedance, Heyi Chen, Siyan Chen, Xin Chen 외 arxiv

Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation.…

Reinforcement LearningVideo Generation

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

2026-06-23 · Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi 외 arxiv

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. Wan-Streamer seamlessly models language, …

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing

2024-01-22 · Xianghu Yue, Xiaohai Tian, Lu Lu, Malu Zhang 외

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…

AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

2026-05-17 · Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang 외 arxiv

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous pr…

Video Generation