paper-with-me

홈 › Papers

Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance Assessments

2023-04-21 · Nirmal Roy, Agathe Balayn, David Maxwell, Claudia Hauff

The creation of relevance assessments by human assessors (often nowadays crowdworkers) is a vital step when building IR test collections. Prior works have investigated assessor quality & behaviour, though into the impact of a document's presentation modality on assessor efficiency and effectiveness. Given the rise of voice-based interfaces, we investigate whether it is feasible for assessors to judge the relevance of text documents via a voice-based interface. We ran a user study (n = 49) on a crowdsourcing platform where participants judged the relevance of short and long documents sampled from the TREC Deep Learning corpus-presented to them either in the text or voice modality. We found that: (i) participants are equally accurate in their judgements across both the text and voice modality; (ii) with increased document length it takes participants significantly longer (for documents of length > 120 words it takes almost twice as much time) to make relevance judgements in the voice condition; and (iii) the ability of assessors to ignore stimuli that are not relevant (i.e., inhibition) impacts the assessment quality in the voice modality-assessors with higher inhibition are significantly more accurate than those with lower inhibition. Our results indicate that we can reliably leverage the voice modality as a means to effectively collect relevance labels from crowdworkers.

📄 PDF Abstract BibTeX arXiv:2304.10881

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

The Munich Biovoice Corpus: Effects of Physical Exercising, Heart Rate, and Skin Conductance on Human Speech Production

2014-05-01 · LREC 2014 5 · Bj{\"o}rn Schuller, Felix Friedmann, Florian Eyben

We introduce a spoken language resource for the analysis of impact that physical exercising has on human speech production. In particular, the database provides heart rate and skin conductance measurement information alo…

Binary ClassificationHeart rate estimation

My lips are concealed: Audio-visual speech enhancement through obstructions

2019-07-11 · Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily…

Speech Enhancement

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

2026-06-04 · Erfan Loweimi, Mengjie Qian, Kate Knill, Guanfeng Wu 외 arxiv

When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or…

Person Retrieval

Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances

2023-06-16 · Huang Xie, Khazar Khorrami, Okko Räsänen, Tuomas Virtanen

This paper explores grading text-based audio retrieval relevances with crowdsourcing assessments. Given a free-form text (e.g., a caption) as a query, crowdworkers are asked to grade audio clips using numeric scores (bet…

Audio captioningContrastive LearningRetrieval

Using Mobile Data and Deep Models to Assess Auditory Verbal Hallucinations

2023-04-20 · Shayan Mirjafari, Subigya Nepal, Weichen Wang, Andrew T. Campbell

Hallucination is an apparent perception in the absence of real external sensory stimuli. An auditory hallucination is a perception of hearing sounds that are not real. A common form of auditory hallucination is hearing v…

HallucinationTransfer Learning