paper-with-me

Papers

Cem Mil Podcasts: A Spoken Portuguese Document Corpus For Multi-modal, Multi-lingual and Multi-Dialect Information Access Research

2022-09-23 · Ekaterina Garmash, Edgar Tanaka, Ann Clifton, Joana Correia, Sharmistha Jat, Winstead Zhu, Rosie Jones, Jussi Karlgren

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the distribution over Brazilian and Portuguese dialects. We give results from experiments on multi-lingual summarization, showing that summarizing podcast transcripts can be performed well by a system supporting both English and Portuguese. We also show experiments on Portuguese podcast genre classification using text metadata. Combining this collection with previously released English-language collection opens up the potential for multi-modal, multi-lingual and multi-dialect podcast information access research.

📄 PDF Abstract BibTeX arXiv:2209.11871

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveGenre classification

Similar Papers 제목 키워드 기반

100,000 Podcasts: A Spoken English Document Corpus

2020-12-01 · COLING 2020 8 · Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu 외

Podcasts are a large and growing repository of spoken audio. As an audio format, podcasts are more varied in style and production type than broadcast news, contain more genres than typically studied in video data, and ar…

3D Facial Landmark LocalizationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Facial Expression Recognition (FER)+5

The COPLE2 corpus: a learner corpus for Portuguese

2016-05-01 · LREC 2016 5 · Am{\'a}lia Mendes, S Antunes, ra, Maarten Janssen 외

We present the COPLE2 corpus, a learner corpus of Portuguese that includes written and spoken texts produced by learners of Portuguese as a second or foreign language. The corpus includes at the moment a total of 182,474…

LemmatizationPOS

It's What You Say and How You Say It: Exploring Textual and Audio Features for Podcast Data

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Podcasts are relatively new media in the form of spoken documents or conversations with a wide range of topics, genres, and styles. With a massive increase in the number of podcasts and their listener base, it is benefic…

TAG

Towards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique

2021-04-12 · Felermino D. M. A. Ali, Andrew Caines, Jaimito L. A. Malavi

Major advancement in the performance of machine translation models has been made possible in part thanks to the availability of large-scale parallel corpora. But for most languages in the world, the existence of such cor…

Machine TranslationSentenceTranslation

Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers

2024-04-10 · Mirelle Bueno, Eduardo Seiti de Oliveira, Rodrigo Nogueira, Roberto A. Lotufo 외

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian…

Information RetrievalRetrieval