paper-with-me

홈 › Papers

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

2025-08-05 · Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Anhong Guo arxiv

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However, the potential benefits and challenges of such capabilities to support diverse real-world assistive tasks remain unclear. In this paper, we present findings from an exploratory study with eight BVI participants. Participants used ChatGPT's Advanced Voice with Video, a state-of-the-art live video AI released in late 2024, in various real-world scenarios, from locating objects to recognizing visual landmarks, across unfamiliar indoor and outdoor environments. Our findings indicate that current live video AI effectively provides guidance and answers for static visual scenes but falls short in delivering essential live descriptions required in dynamic situations. Despite inaccuracies in spatial and distance information, participants leveraged the provided visual information to supplement their mobility strategies. Although the system was perceived as human-like due to high-quality voice interactions, assumptions about users' visual abilities, hallucinations, generic responses, and a tendency towards sycophancy led to confusion, distrust, and potential risks for BVI users. Based on the results, we discuss implications for assistive video AI agents, including incorporating additional sensing capabilities for real-world use, determining appropriate intervention timing beyond turn-taking interactions, and addressing ecological and safety concerns.

📄 PDF Abstract BibTeX arXiv:2508.03651

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

2023-11-22 · Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed 외

Extending image-based Large Multimodal Models (LMMs) to videos is challenging due to the inherent complexity of video data. The recent approaches extending image-based LMMs to videos either lack the grounding capabilitie…

BenchmarkingPhrase GroundingQuestion AnsweringSpatio-Temporal Video Grounding+2

Can ChatGPT capture swearing nuances? Evidence from translating Arabic oaths

2024-12-03 · Mohammed Q. Shormani

This study sets out to answer one major question: Can ChatGPT capture swearing nuances? It presents an empirical study on the ability of ChatGPT to translate Arabic oath expressions into English. 30 Arabic oath expressio…

Translation

Seeing ChatGPT Through Students' Eyes: An Analysis of TikTok Data

2023-03-09 · Anna-Carolina Haensch, Sarah Ball, Markus Herklotz, Frauke Kreuter

Advanced large language models like ChatGPT have gained considerable attention recently, including among students. However, while the debate on ChatGPT in academia is making waves, more understanding is needed among lect…

ChatGPT and U(X): A Rapid Review on Measuring the User Experience

2025-03-20 · Katie Seaborn

ChatGPT, powered by a large language model (LLM), has revolutionized everyday human-computer interaction (HCI) since its 2022 release. While now used by millions around the world, a coherent pathway for evaluating the us…

Language ModelingLanguage ModellingLarge Language Model

Evade ChatGPT Detectors via A Single Space

2023-07-05 · Shuyang Cai, Wanyun Cui

ChatGPT brings revolutionary social value but also raises concerns about the misuse of AI-generated text. Consequently, an important question is how to detect whether texts are generated by ChatGPT or by human. Existing …

Language ModelingLanguage Modelling