paper-with-me

Papers

Augmented Conversation with Embedded Speech-Driven On-the-Fly Referencing in AR

2024-05-28 · Shivesh Jadon, Mehrad Faridan, Edward Mah, Rajan Vaish, Wesley Willett, Ryo Suzuki

This paper introduces the concept of augmented conversation, which aims to support co-located in-person conversations via embedded speech-driven on-the-fly referencing in augmented reality (AR). Today computing technologies like smartphones allow quick access to a variety of references during the conversation. However, these tools often create distractions, reducing eye contact and forcing users to focus their attention on phone screens and manually enter keywords to access relevant information. In contrast, AR-based on-the-fly referencing provides relevant visual references in real-time, based on keywords extracted automatically from the spoken conversation. By embedding these visual references in AR around the conversation partner, augmented conversation reduces distraction and friction, allowing users to maintain eye contact and supporting more natural social interactions. To demonstrate this concept, we developed \system, a Hololens-based interface that leverages real-time speech recognition, natural language processing and gaze-based interactions for on-the-fly embedded visual referencing. In this paper, we explore the design space of visual referencing for conversations, and describe our our implementation -- building on seven design guidelines identified through a user-centered design process. An initial user study confirms that our system decreases distraction and friction in conversations compared to smartphone searches, while providing highly useful and relevant information.

📄 PDF Abstract BibTeX arXiv:2405.18537

Code (0)

등록된 구현이 없습니다.

Tasks

Frictionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models

2025-01-20 · Jingwei Yi, Junhao Yin, Ju Xu, Peng Bao 외

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in understanding multimodal inputs and have been widely integrated into Retrieval-Augmented Generation (RAG) based conversational systems. While cur…

RAGRetrievalRetrieval-augmented Generation

RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling

2022-08-12 · Jian Liao, Adnan Karim, Shivesh Jadon, Rubaiat Habib Kazi 외

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive sto…

Video Editing

GASCOM: Graph-based Attentive Semantic Context Modeling for Online Conversation Understanding

2023-10-21 · Vibhor Agarwal, Yu Chen, Nishanth Sastry

Online conversation understanding is an important yet challenging NLP problem which has many useful applications (e.g., hate speech detection). However, online conversations typically unfold over a series of posts and re…

Graph AttentionHate Speech Detection

Improving YOLOv8 with Scattering Transform and Attention for Maritime Awareness

2023-10-19 · International Symposium on Image and Signal Processing and Analysis (ISPA) 2023 10 · Borja Carrillo-Perez, Angel Bueno Rodriguez, Sarah Barnes, Maurice Stephan

Ship recognition and georeferencing using monitoring cameras are crucial to many applications in maritime situational awareness. Although deep learning algorithms are available for ship recognition tasks, there is a need…

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation

2022-07-29 · Giulia Comini, Goeric Huybrechts, Manuel Sam Ribeiro, Adam Gabrys 외

The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming. To overcome these issues, we demonstrate how to build low-resource, neural text-to-speech (…

Data Augmentationtext-to-speechText to SpeechVoice Conversion