paper-with-me

홈 › Papers

Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

2022-02-04 · André F. A. Paschoal, Paulo Pirozelli, Valdinei Freire, Karina V. Delgado, Sarajane M. Peres, Marcos M. José, Flávio Nakasato, André S. Oliveira, Anarosa A. F. Brandão, Anna H. R. Costa, Fabio G. Cozman

Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.

📄 PDF Abstract BibTeX arXiv:2202.02398

Code (1)

C4AI/Pira 공식 구현

Tasks

Information RetrievalMachine TranslationQuestion AnsweringRetrieval

Similar Papers 제목 키워드 기반

Cross-Lingual Named Entity Recognition via FastAlign: a Case Study

2021-07-01 · TRITON 2021 7 · Ali Hatami, Ruslan Mitkov, Gloria Corpas Pastor

Named Entity Recognition is an essential task in natural language processing to detect entities and classify them into predetermined categories. An entity is a meaningful word, or phrase that refers to proper nouns. Name…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

BiMediX: Bilingual Medical Mixture of Experts LLM

2024-02-20 · Sara Pieri, Sahal Shaji Mullappilly, Fahad Shahbaz Khan, Rao Muhammad Anwer 외

In this paper, we introduce BiMediX, the first bilingual medical mixture of experts LLM designed for seamless interaction in both English and Arabic. Our model facilitates a wide range of medical interactions in English …

Mixture-of-ExpertsMultiple-choiceOpen-Ended Question AnsweringQuestion Answering

Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

2022-10-04 · COLING 2022 10 · Priyanka Sen, Alham Fikri Aji, Amir Saffari

We introduce Mintaka, a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models. Mintaka is composed of 20,000 question-answer pairs collected in English, annotated…

Question Answering

Event-QA: A Dataset for Event-Centric Question Answering over Knowledge Graphs

2020-04-24 · Tarcísio Souza Costa, Simon Gottschalk, Elena Demidova

Semantic Question Answering (QA) is a crucial technology to facilitate intuitive user access to semantic information stored in knowledge graphs. Whereas most of the existing QA systems and datasets focus on entity-centri…

Knowledge GraphsQuestion Answering

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

2026-04-15 · Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires, Celio Larcher 외 arxiv

Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesize complete answers. While recent benchm…