paper-with-me

홈 › Papers

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

2026-06-19 · Jiehui Huang, Yuechen Zhang, Bin Xia, Jiahao Wang, Xu He, Zhenchao Tang, Meng Chu, Xin Tao, Pengfei Wan, Jiaya Jia arxiv

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-end over fixed-length sequences and cannot scale, generate shot-by-shot with memory banks that grow linearly, or orchestrate pretrained generators under an LLM planner without a multi-shot-aware backbone. We present UnityShots, a memory-driven multi-shot audio-video generation system built on LTX-2.3, trained on annotated cinematic and music-video shots. The video stream maintains two fixed-size slots, a long-term memory (LTM) slot anchored to the opening shot and a short-term memory (STM) slot holding the immediately preceding tail, both updated at every cut by a boundary-conditioned gate that fuses visual cut probability and beat-tracker signals. The audio stream injects a reference speaker token at every shot to preserve vocal timbre without a sliding audio bank. A discrete cut-type prior, learned through AdaLN, becomes an inference-time control knob over transition strength. We release a benchmark of $200$ multi-cultural multi-shot sequences spanning six ethnic regions and ten or more languages, with per-shot reference identities, reference audio, and per-boundary transition labels. Evaluated across I2V, T2V, and R2V conditioning modes, UnityShots leads open-source baselines on every cross-shot coherence metric and matches the strongest closed-source system on the multi-shot axes.

📄 PDF Abstract BibTeX arXiv:2606.21661

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation

S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis

2024-08-18 · Dongze Li, Kang Zhao, Wei Wang, Yifeng Ma 외

Talking head synthesis is a practical technique with wide applications. Current Neural Radiance Field (NeRF) based approaches have shown their superiority on driving one-shot talking heads with videos or signals regresse…

NeRF

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

2026-05-06 · Leying Zhang, Bowen Shi, Haibin Wu, Bach Viet Do 외 arxiv

The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain…

Zero-shot GeneralizationDomain GeneralizationInstruction Following

One Shot Audio to Animated Video Generation

2021-02-19 · Neeraj Kumar, Srishti Goel, Ankur Narang, Brejesh lall 외

We consider the challenging problem of audio to animated video generation. We propose a novel method OneShotAu2AV to generate an animated video of arbitrary length using an audio clip and a single unseen image of a perso…

Video Generation

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

2021-12-06 · Suzhen Wang, Lincheng Li, Yu Ding, Xin Yu

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those metho…

Face GenerationTalking Face Generation