paper-with-me

Papers

RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling

2022-08-12 · Jian Liao, Adnan Karim, Shivesh Jadon, Rubaiat Habib Kazi, Ryo Suzuki

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However, existing tools for live presentations often lack interactivity and improvisation, while creating such effects in video editing tools require significant time and expertise. RealityTalk enables users to create live augmented presentations with real-time speech-driven interactions. The user can interactively prompt, move, and manipulate graphical elements through real-time speech and supporting modalities. Based on our analysis of 177 existing video-edited augmented presentations, we propose a novel set of interaction techniques and then incorporated them into RealityTalk. We evaluate our tool from a presenter's perspective to demonstrate the effectiveness of our system.

📄 PDF Abstract BibTeX arXiv:2208.06350

Code (0)

등록된 구현이 없습니다.

Tasks

Video Editing

Similar Papers 제목 키워드 기반

Augmented Conversation with Embedded Speech-Driven On-the-Fly Referencing in AR

2024-05-28 · Shivesh Jadon, Mehrad Faridan, Edward Mah, Rajan Vaish 외

This paper introduces the concept of augmented conversation, which aims to support co-located in-person conversations via embedded speech-driven on-the-fly referencing in augmented reality (AR). Today computing technolog…

Frictionspeech-recognitionSpeech Recognition

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

2024-08-03 · Yihong Lin, Zhaoxin Fan, Xianjia Wu, Lingyu Xiong 외

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have ach…

DiversityTalking Head Generation

KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation

2025-09-24 · Tianle Lyu, Junchuan Zhao, Ye Wang arxiv

Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for talking-face synthesis. However, most existing works treat speech features as a m…

Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator

2024-03-25 · Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka

A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-…

Data AugmentationGenerative Adversarial NetworkSpeech Synthesis

Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection

2022-05-09 · NAACL 2022 7 · Indira Sen, Mattia Samory, Claudia Wagner, Isabelle Augenstein

Counterfactually Augmented Data (CAD) aims to improve out-of-domain generalizability, an indicator of model robustness. The improvement is credited with promoting core features of the construct over spurious artifacts th…

Hate Speech Detection