Cem Mil Podcasts: A Spoken Portuguese Document Corpus For Multi-modal, Multi-lingual and Multi-Dialect Information Access Research
In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the distribution over Brazilian and Portuguese dialects. We give results from experiments on multi-lingual summarization, showing that summarizing podcast transcripts can be performed well by a system supporting both English and Portuguese. We also show experiments on Portuguese podcast genre classification using text metadata. Combining this collection with previously released English-language collection opens up the potential for multi-modal, multi-lingual and multi-dialect podcast information access research.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveGenre classificationSimilar Papers 제목 키워드 기반
100,000 Podcasts: A Spoken English Document Corpus
Podcasts are a large and growing repository of spoken audio. As an audio format, podcasts are more varied in style and production type than broadcast news, contain more genres than typically studied in video data, and ar…
3D Facial Landmark LocalizationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Facial Expression Recognition (FER)+5The COPLE2 corpus: a learner corpus for Portuguese
We present the COPLE2 corpus, a learner corpus of Portuguese that includes written and spoken texts produced by learners of Portuguese as a second or foreign language. The corpus includes at the moment a total of 182,474…
LemmatizationPOSIt's What You Say and How You Say It: Exploring Textual and Audio Features for Podcast Data
Podcasts are relatively new media in the form of spoken documents or conversations with a wide range of topics, genres, and styles. With a massive increase in the number of podcasts and their listener base, it is benefic…
TAGTowards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique
Major advancement in the performance of machine translation models has been made possible in part thanks to the availability of large-scale parallel corpora. But for most languages in the world, the existence of such cor…
Machine TranslationSentenceTranslationQuati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers
Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian…
Information RetrievalRetrieval