paper-with-me

Papers

Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation

2026-08-07 · Lunjie Zhu, Xingtong Ge, Fangyu Lin, Yi Zhang, Zhening Liu, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang arxiv

Joint audio-video generative models serve as foundation for immersive and interactive digital-human generation. Nevertheless, most existing models rely on bidirectional attention and multi-step denoising and can generate only short clips, making them unsuitable for real-time interaction over extended durations. We present Omni-LiveAvatar, the first framework for minute-level, real-time streaming joint audio-video avatar generation. Specifically, we propose (1) a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary stabilization mechanisms; (2) a synchronized audio-video long-short-term memory that preserves global consistency under a bounded memory budget; and (3) a hierarchical rolling prompt planning strategy that enables coherent semantic evolution and seamless prompt transitions. Extensive experiments show that Omni-LiveAvatar generates high-quality, synchronized minute-level avatars in real time. In terms of speed, it achieves a 33$\times$ generation speedup over its teacher, LTX-2, on a single NVIDIA H200 GPU; in terms of generation quality, it outperforms accelerated baselines across visual quality, audio quality, cross-modal synchronization, and human fidelity. Our code is available at https://github.com/Aoko955/Omni-LiveAvatar.

📄 PDF Abstract BibTeX arXiv:2608.13602

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

2026-03-19 · Keda Tao, Yuhua Zheng, Jia Xu, Wenjie Du 외 arxiv

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips rangi…

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

2026-05-21 · Haichen He, Jiayi Zhou, Sifeng Shang, Yihan Hu 외 arxiv

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitiv…

An Accelerated Pipeline for Multi-label Renal Pathology Image Segmentation at the Whole Slide Image Level

2023-05-23 · Haoju Leng, Ruining Deng, Zuhayr Asad, R. Michael Womick 외

Deep-learning techniques have been used widely to alleviate the labour-intensive and time-consuming manual annotation required for pixel-level tissue characterization. Our previous study introduced an efficient single dy…

GPUImage SegmentationSegmentationSemantic Segmentation+1

Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion

2025-03-11 · Nico Bohlinger, Jonathan Kinzel, Daniel Palenicek, Lukasz Antczak 외

On-robot Reinforcement Learning is a promising approach to train embodiment-aware policies for legged robots. However, the computational constraints of real-time learning on robots pose a significant challenge. We presen…

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

2026-07-04 · Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han 외 hf

Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact…