paper-with-me

홈 › Papers

Navigating Speech Recording Collections with AI-Generated Illustrations

2025-07-05 · Sirina Håland, Trond Karlsen Strøm, Petra Galuščáková arxiv

Although the amount of available spoken content is steadily increasing, extracting information and knowledge from speech recordings remains challenging. Beyond enhancing traditional information retrieval methods such as speech search and keyword spotting, novel approaches for navigating and searching spoken content need to be explored and developed. In this paper, we propose a novel navigational method for speech archives that leverages recent advances in language and multimodal generative models. We demonstrate our approach with a Web application that organizes data into a structured format using interactive mind maps and image generation tools. The system is implemented using the TED-LIUM~3 dataset, which comprises over 2,000 speech transcripts and audio files of TED Talks. Initial user tests using a System Usability Scale (SUS) questionnaire indicate the application's potential to simplify the exploration of large speech collections.

📄 PDF Abstract BibTeX arXiv:2507.04182

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalImage GenerationKeyword Spotting

Similar Papers 제목 키워드 기반

Manual Speech Synthesis Data Acquisition - From Script Design to Recording Speech

2020-05-01 · LREC 2020 5 · Atli Sigurgeirsson, Gunnar {\"O}rn{\'o}lfsson, J{\'o}n Gu{\dh}nason

Atli {\TH}{\'o}r Sigurgeirsson, atlithors@ru.is, Reykjavik University Gunnar Thor {\"O}rn{\'o}lfsson, gunnarthor@hi.is, {\'A}rni Magn{\'u}sson institute of Icelandic studies Dr. J{\'o}n Gu{\dh}nason, jg@ru.is In this pap…

Speech Synthesis

UltraSuite: A Repository of Ultrasound and Acoustic Data from Child Speech Therapy Sessions

2019-07-01 · Aciel Eshky, Manuel Sam Ribeiro, Joanne Cleland, Korin Richmond 외

We introduce UltraSuite, a curated repository of ultrasound and acoustic data, collected from recordings of child speech therapy sessions. This release includes three data collections, one from typically developing child…

Speech Resources in the Tamasheq Language

2022-01-13 · LREC 2022 6 · Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche 외

In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist…

Translation

Endangered Language Documentation: Bootstrapping a Chatino Speech Corpus, Forced Aligner, ASR

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Hilaria Cruz

This project approaches the problem of language documentation and revitalization from a rather untraditional angle. To improve and facilitate language documentation of endangered languages, we attempt to use corpus lingu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

2023-06-07 · Claytone Sikasote, Kalinda Siaminwe, Stanly Mwape, Bangiwe Zulu 외

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours…

Cross-Lingual Transferspeech-recognitionSpeech RecognitionTransfer Learning