paper-with-me

홈 › Papers

A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls

2026-08-28 · Mirae Kim, Seonghun Jeong, Youngjun Kwak arxiv

Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what executives say, how they say it carries independent and complementary information. To study these dimensions jointly, we introduce DualEvasion, a benchmark for evasion detection across text and audio in earnings call Q&A. The benchmark contains 505 annotated question-answer pairs from 60 earnings calls, each with two independent labels: textual evasion (direct vs. evasive) and vocal cues operationalized as speaker confidence (confident vs. unconfident). Our experiments show that state-of-the-art multimodal models struggle to detect vocal confidence, particularly on unconfident responses. Our analysis suggests these models interpret acoustic cues in isolation rather than relative to each speaker's baseline. Providing speaker-level references yields modest improvements, but a substantial gap with human performance remains.

📄 PDF Abstract BibTeX arXiv:2608.28040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DODGE: Ontology-Aware Risk Assessment via Object-Oriented Disruption Graphs

2024-12-18 · Stefano M. Nicoletti, E. Moritz Hahn, Mattia Fumagalli, Giancarlo Guizzardi 외

When considering risky events or actions, we must not downplay the role of involved objects: a charged battery in our phone averts the risk of being stranded in the desert after a flat tyre, and a functional firewall mit…

Adaptively Meshed Video Stabilization

2020-06-14 · Minda Zhao, Qiang Ling

Video stabilization is essential for improving visual quality of shaky videos. The current video stabilization methods usually take feature trajectories in the background to estimate one global transformation matrix or s…

BlockingMotion EstimationVideo Stabilization

ShakyPrepend: A Multi-Group Learner with Improved Sample Complexity

2026-03-07 · Lujing Zhang, Daniel Hsu, Sivaraman Balakrishnan arxiv

Multi-group learning is a learning task that focuses on controlling predictors' conditional losses over specified subgroups. We propose ShakyPrepend, a method that leverages tools inspired by differential privacy to obta…

Dodgersort: Uncertainty-Aware VLM-Guided Human-in-the-Loop Pairwise Ranking

2026-03-21 · Yujin Park, Haejun Chung, Ikbeom Jang arxiv

Pairwise comparison labeling is emerging as it yields higher inter-rater reliability than conventional classification labeling, but exhaustive comparisons require quadratic cost. We propose Dodgersort, which leverages CL…

MLNET: An Adaptive Multiple Receptive-field Attention Neural Network for Voice Activity Detection

2020-08-13 · Zhenpeng Zheng, Jianzong Wang, Ning Cheng, Jian Luo 외

Voice activity detection (VAD) makes a distinction between speech and non-speech and its performance is of crucial importance for speech based services. Recently, deep neural network (DNN)-based VADs have achieved better…

Action DetectionActivity Detection