paper-with-me

홈 › Papers

Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents

2025-05-20 · Takao Fujii, Katie Seaborn, Madeleine Steeds, Jun Kato

Conversational agents that mimic people have raised questions about the ethics of anthropomorphizing machines with human social identity cues. Critics have also questioned assumptions of identity neutrality in humanlike agents. Recent work has revealed that intersectional Japanese pronouns can elicit complex and sometimes evasive impressions of agent identity. Yet, the role of other "neutral" non-pronominal self-referents (NPSR) and voice as a socially expressive medium remains unexplored. In a crowdsourcing study, Japanese participants (N = 204) evaluated three ChatGPT voices (Juniper, Breeze, and Ember) using seven self-referents. We found strong evidence of voice gendering alongside the potential of intersectional self-referents to evade gendering, i.e., ambiguity through neutrality and elusiveness. Notably, perceptions of age and formality intersected with gendering as per sociolinguistic theories, especially boku and watakushi. This work provides a nuanced take on agent identity perceptions and champions intersectional and culturally-sensitive work on voice agents.

📄 PDF Abstract BibTeX arXiv:2506.01998

Code (0)

등록된 구현이 없습니다.

Tasks

Ethics

Similar Papers 제목 키워드 기반

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

2026-03-19 · Shree Harsha Bokkahalli Satish, Maria Teleki, Christoph Minixhofer, Ondrej Klejch 외 arxiv

SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations often miss how such bias manifests in end-to-end speech interactions and h…

Voice Conversion

The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs

2026-03-15 · Shree Harsha Bokkahalli Satish, Christoph Minixhofer, Maria Teleki, James Caverlee 외 arxiv

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent v…

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

2025-11-10 · Hyeryun Park, Byung Mo Gu, Jun Hee Lee, Byeong Hyeon Choi 외 arxiv

In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, …

Discovering the Italian literature: interactive access to audio indexed text resources

2014-05-01 · LREC 2014 5 · Vincenzo Galat{\`a}, Alberto Benin, Piero Cosi, Giuseppe Riccardo Leone 외

In this paper we present a web interface to study Italian through the access to read Italian literature. The system allows to browse the content, search for specific words and listen to the correct pronunciation produced…

Cultural Vocal Bursts Intensity PredictionSentencetext-to-speechText to Speech

``Voices of the Great War'': A Richly Annotated Corpus of Italian Texts on the First World War

2020-05-01 · LREC 2020 5 · Federico Boschetti, Irene De Felice, Stefano Dei Rossi, Felice Dell{'}Orletta 외

{``}Voices of the Great War{''} is the first large corpus of Italian historical texts dating back to the period of First World War. This corpus differs from other existing resources in several respects. First, from the l…