paper-with-me

홈 › Papers

Linguistically Augmented Audio Speech Data (LinguAS)

2026-06-08 · Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja arxiv

Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve. Yet, most detection models are trained to make inference on frame-level audio features alone without leveraging valuable linguistic cues at larger timescales. To address this gap, we present Linguistically Augmented Audio Speech Data (LinguAS), a dataset of genuine and deepfaked audio samples annotated with five strategically-chosen, Expert-Defined Linguistic Features (EDLFs) that occur frequently in spoken English and are characteristic of natural human speech. LinguAS contains over 800 audio samples, each of which are annotated with EDLFs. The dataset has a balanced number of four spoofed audio attack types and a proportionate number of genuine speech samples. We also include metadata on speaker gender and the generator/source for each spoofed audio sample, offering more granularity for model training. We found that models trained on data augmented with EDLFs had improved model performance significantly beyond the ASVspoof 2021 deep learning baselines and SSL models like HuBert and XLSR. LinguAS's augmented linguistic, gender, and generator metadata provide audio deepfake researchers with a dataset that emphasizes real human language traits to improve model inference of faked speech. Data and code are publicly available.

📄 PDF Abstract BibTeX arXiv:2606.10246

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

L2 proficiency assessment using self-supervised speech representations

2022-11-16 · Stefano Bannò, Kate M. Knill, Marco Matassoni, Vyas Raina 외

There has been a growing demand for automated spoken language assessment systems in recent years. A standard pipeline for this process is to start with a speech recognition system and derive features, either hand-crafted…

speech-recognitionSpeech Recognition

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

2026-06-03 · Eugene Kwek, Feng Liu, Rui Zhang, Wenpeng Yin arxiv

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token effic…

Voice Conversion

LinguaSim: Interactive Multi-Vehicle Testing Scenario Generation via Natural Language Instruction Based on Large Language Models

2025-10-09 · Qingyuan Shi, Qingwen Meng, Hao Cheng, Qing Xu 외 arxiv

The generation of testing and training scenarios for autonomous vehicles has drawn significant attention. While Large Language Models (LLMs) have enabled new scenario generation methods, current methods struggle to balan…

Autonomous VehiclesAutonomous Driving

Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

2022-06-30 · Hyeon-Kyeong Shin, Hyewon Han, Doyeon Kim, Soo-Whan Chung 외

In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike previous approaches requiring speech keyword…

Keyword Spotting

When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper

2026-03-05 · Akif Islam, Raufun Nahar, Md. Ekramul Hamid arxiv

Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigo…

Speech RecognitionSpeech Enhancement