paper-with-me

홈 › Papers

PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents

2024-10-10 · Leandro Carísio Fernandes, Guilherme Zeferino Rodrigues Dobins, Roberto Lotufo, Jayr Alencar Pereira

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with news articles and structured summaries containing the individuals participating in the hearing and their statements or opinions. The dataset supports the development and evaluation of long document summarization systems in Portuguese. Our contributions include the dataset, a hybrid summarization system to establish a baseline for future studies, and a discussion on evaluation metrics for summarization involving large language models, addressing the challenge of hallucination in the generated summaries. As a result of this discussion, the dataset also provides annotated data that can be used in Natural Language Inference tasks in Portuguese.

📄 PDF Abstract BibTeX arXiv:2410.07495

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDocument SummarizationHallucinationNatural Language Inference

Similar Papers 제목 키워드 기반

Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence

2024-09-06 · Bruno S. Masiero, Leticia R. Borges, Harvey Dillon, Maria Francisca Colella-Santos

The Listening in Spatialized Noise Sentences (LiSN-S) is a test to evaluate auditory spatial processing currently only available in the English language. It produces a three-dimensional auditory environment under headpho…

Sentence

Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing

2020-12-01 · COLING 2020 8 · Felipe Almeida Costa, Thiago castro Ferreira, Adriana Pagano, Wagner Meira

This paper introduces the first corpus for Automatic Post-Editing of English and a low-resource language, Brazilian Portuguese. The source English texts were extracted from the WebNLG corpus and automatically translated …

Automatic Post-EditingMachine TranslationTranslation

TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese

2020-05-11 · Edresson Casanova, Arnaldo Candido Junior, Christopher Shulby, Frederico Santos de Oliveira 외

Speech provides a natural way for human-computer interaction. In particular, speech synthesis systems are popular in different applications, such as personal assistants, GPS applications, screen readers and accessibility…

DenoisingSpeech SynthesisTransfer Learning

Image captioning for Brazilian Portuguese using GRIT model

2024-02-07 · Rafael Silva de Alencar, William Alberto Cruz Castañeda, Marcellus Amadeus

This work presents the early development of a model of image captioning for the Brazilian Portuguese language. We used the GRIT (Grid - and Region-based Image captioning Transformer) model to accomplish this work. GRIT i…

Image Captioningmodel

Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers

2024-04-10 · Mirelle Bueno, Eduardo Seiti de Oliveira, Rodrigo Nogueira, Roberto A. Lotufo 외

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian…

Information RetrievalRetrieval