paper-with-me

홈 › Papers

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

2025-07-14 · Youliang Zhang, Zhaoyang Li, Duomin Wang, Jiahe Zhang, Deyu Zhou, Zixin Yin, Xili Dai, Gang Yu, Xiu Li arxiv

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the next major challenge: audio-visual dyadic interactive virtual human. To facilitate research in this emerging area, we present SpeakerVid-5M dataset, the first large-scale, high-quality dataset designed for audio-visual dyadic interactive virtual human generation. Totaling over 8,743 hours, SpeakerVid-5M contains more than 5.2 million video clips of human portraits. It covers diverse scales and interaction types, including monadic talking, listening, and dyadic conversations. Crucially, the dataset is structured along two key dimensions: interaction type and data quality. First, it is categorized into four types (dialogue branch, single branch, listening branch and multi-turn branch) based on the interaction scenario. Second, it is stratified into a large-scale pre-training subset and a curated, high-quality subset for Supervised Fine-Tuning (SFT). This dual structure accommodates a wide array of 2D virtual human tasks. In addition, we provide an autoregressive (AR)-based video chat baseline trained on this data, accompanied by a dedicated set of metrics and test data to serve as a benchmark VidChatBench for future work. Both the dataset and the corresponding data processing code will be publicly released. Project page: https://dorniwang.github.io/SpeakerVid-5M/

📄 PDF Abstract BibTeX arXiv:2507.09862

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data Quality Measures and Efficient Evaluation Algorithms for Large-Scale High-Dimensional Data

2021-01-05 · Hyeongmin Cho, Sangkyun Lee

Machine learning has been proven to be effective in various application areas, such as object and speech recognition on mobile systems. Since a critical key to machine learning success is the availability of large traini…

BIG-bench Machine Learningspeech-recognitionSpeech Recognition

OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing

2025-12-08 · Haoyang He, Jie Wang, Jiangning Zhang, Zhucun Xue 외 arxiv

The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introdu…

Image Editing

SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data

2025-11-23 · Sultan Alrashed, Chadi Helwe, Francesco Orabona arxiv

Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool calling. Naive translation can work at the…

Scalable Vision Language Model Training via High Quality Data Curation

2025-01-10 · Hongyuan Dong, Zijian Kang, Weijie Yin, Xiao Liang 외

In this paper, we introduce SAIL-VL (ScAlable Vision Language Model TraIning via High QuaLity Data Curation), an open-source vision language model (VLM) of state-of-the-art (SOTA) performance with 2B parameters. We intro…

Instruction FollowingLanguage ModelingLanguage Modelling

Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

2025-10-01 · Thiziri Nait Saada, Louis Bethune, Michal Klein, David Grangier 외 arxiv

Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quality Filtering (CQF), which trains a binar…