paper-with-me

홈 › Papers

ViDA-MAN: Visual Dialog with Digital Humans

2021-10-26 · Tong Shen, Jiawei Zuo, Fan Shi, Jin Zhang, Liqin Jiang, Meng Chen, Zhengchen Zhang, Wei zhang, Xiaodong He, Tao Mei

We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge.

📄 PDF Abstract BibTeX arXiv:2110.13384

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognitiontext-to-speechText to SpeechVideo GenerationVisual Dialog

Similar Papers 제목 키워드 기반

1000x Faster Camera and Machine Vision with Ordinary Devices

2022-01-23 · Tiejun Huang, Yajing Zheng, Zhaofei Yu, Rui Chen 외

In digital cameras, we find a major limitation: the image and video form inherited from a film camera obstructs it from capturing the rapidly changing photonic world. Here, we present vidar, a bit sequence array where ea…

Deblurringobject-detectionObject Detection

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

2026-06-30 · Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele 외 arxiv

Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sess…

Visual Grounding

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

2025-11-16 · Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li 외 arxiv

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversation…

Dialogue GenerationResponse GenerationSpeech Synthesis

\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation

2026-01-26 · Weiye Zhu, Zekai Zhang, Xiangchen Wang, Hewei Pan 외 arxiv

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings with…

Vision-Language Navigation

Conversational Human Audio-visual Talking Dialogue Generation

2026-07-02 · Junhao Song, Lluis Guasch, Xilin He, Zhongyu Yang 외 arxiv

Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. However, collecting such data is time-consumi…

Dialogue Generation