paper-with-me

홈 › Papers

On Barriers to Archival Audio Processing

2025-07-11 · Peter Sullivan, Muhammad Abdul-Mageed arxiv

In this study, we leverage a unique UNESCO collection of mid-20th century radio recordings to probe the robustness of modern off-the-shelf language identification (LID) and speaker recognition (SR) methods, especially with respect to the impact of multilingual speakers and cross-age recordings. Our findings suggest that LID systems, such as Whisper, are increasingly adept at handling second-language and accented speech. However, speaker embeddings remain a fragile component of speech processing pipelines that is prone to biases related to the channel, age, and language. Issues which will need to be overcome should archives aim to employ SR methods for speaker indexing.

📄 PDF Abstract BibTeX arXiv:2507.08768

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSpeaker Recognition

Similar Papers 제목 키워드 기반

Keyword spotting for audiovisual archival search in Uralic languages

2021-09-01 · ACL (IWCLUL) 2021 9 · Nils Hjortnaes, Niko Partanen, Francis M. Tyers
Keyword Spotting

Automated speech tools for helping communities process restricted-access corpora for language revival efforts

2022-04-15 · ComputEL (ACL) 2022 5 · Nay San, Martijn Bartelds, Tolúlopé Ògúnrèmí, Alison Mount 외

Many archival recordings of speech from endangered languages remain unannotated and inaccessible to community members and language learning programs. One bottleneck is the time-intensive nature of annotation. An even nar…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

AI Blob! LLM-Driven Recontextualization of Italian Television Archives

2025-08-13 · Roberto Balestri arxiv

This paper introduces AI Blob!, an experimental system designed to explore the potential of semantic cataloging and Large Language Models (LLMs) for the retrieval and recontextualization of archival television footage. D…

Speech Recognition

POST: Email Archival, Processing and Flagging Stack for Incident Responders

2024-07-01 · Jeffrey Fairbanks

Phishing is one of the main points of compromise, with email security and awareness being estimated at \$50-100B in 2022. There is great need for email forensics capability to quickly search for malicious content. A nove…

Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)

2025-12-22 · Niclas Griesshaber, Jochen Streb arxiv

We leverage multimodal large language models (LLMs) to construct a dataset of 306,070 German patents (1877-1918) from 9,562 archival image scans using our LLM-based pipeline powered by Gemini-2.5-Pro and Gemini-2.5-Flash…