Navigating Speech Recording Collections with AI-Generated Illustrations
Although the amount of available spoken content is steadily increasing, extracting information and knowledge from speech recordings remains challenging. Beyond enhancing traditional information retrieval methods such as speech search and keyword spotting, novel approaches for navigating and searching spoken content need to be explored and developed. In this paper, we propose a novel navigational method for speech archives that leverages recent advances in language and multimodal generative models. We demonstrate our approach with a Web application that organizes data into a structured format using interactive mind maps and image generation tools. The system is implemented using the TED-LIUM~3 dataset, which comprises over 2,000 speech transcripts and audio files of TED Talks. Initial user tests using a System Usability Scale (SUS) questionnaire indicate the application's potential to simplify the exploration of large speech collections.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalImage GenerationKeyword SpottingSimilar Papers 제목 키워드 기반
Manual Speech Synthesis Data Acquisition - From Script Design to Recording Speech
Atli {\TH}{\'o}r Sigurgeirsson, atlithors@ru.is, Reykjavik University Gunnar Thor {\"O}rn{\'o}lfsson, gunnarthor@hi.is, {\'A}rni Magn{\'u}sson institute of Icelandic studies Dr. J{\'o}n Gu{\dh}nason, jg@ru.is In this pap…
Speech SynthesisUltraSuite: A Repository of Ultrasound and Acoustic Data from Child Speech Therapy Sessions
We introduce UltraSuite, a curated repository of ultrasound and acoustic data, collected from recordings of child speech therapy sessions. This release includes three data collections, one from typically developing child…
Speech Resources in the Tamasheq Language
In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist…
TranslationEndangered Language Documentation: Bootstrapping a Chatino Speech Corpus, Forced Aligner, ASR
This project approaches the problem of language documentation and revitalization from a rather untraditional angle. To improve and facilitate language documentation of endangered languages, we attempt to use corpus lingu…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionZambezi Voice: A Multilingual Speech Corpus for Zambian Languages
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours…
Cross-Lingual Transferspeech-recognitionSpeech RecognitionTransfer Learning