paper-with-me

Papers

BiPaR: A Bilingual Parallel Dataset for Multilingual and Cross-lingual Reading Comprehension on Novels

2019-10-11 · IJCNLP 2019 11 · Yimin Jing, Deyi Xiong, Yan Zhen

This paper presents BiPaR, a bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support multilingual and cross-lingual reading comprehension. The biggest difference between BiPaR and existing reading comprehension datasets is that each triple (Passage, Question, Answer) in BiPaR is written parallelly in two languages. We collect 3,667 bilingual parallel paragraphs from Chinese and English novels, from which we construct 14,668 parallel question-answer pairs via crowdsourced workers following a strict quality control procedure. We analyze BiPaR in depth and find that BiPaR offers good diversification in prefixes of questions, answer types and relationships between questions and passages. We also observe that answering questions of novels requires reading comprehension skills of coreference resolution, multi-sentence reasoning, and understanding of implicit causality, etc. With BiPaR, we build monolingual, multilingual, and cross-lingual MRC baseline models. Even for the relatively simple monolingual MRC on this dataset, experiments show that a strong BERT baseline is over 30 points behind human in terms of both EM and F1 score, indicating that BiPaR provides a challenging testbed for monolingual, multilingual and cross-lingual MRC on novels. The dataset is available at https://multinlp.github.io/BiPaR/.

📄 PDF Abstract BibTeX arXiv:1910.05040

Code (1)

sharejing/BiPaR

Tasks

coreference-resolutionCoreference ResolutionMachine Reading ComprehensionReading ComprehensionSentence

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Advancing Multilingual Pre-training: TRIP Triangular Document-level Pre-training for Multilingual Language Models

2022-12-15 · Hongyuan Lu, Haoyang Huang, Shuming Ma, Dongdong Zhang 외

Despite the success of multilingual sequence-to-sequence pre-training, most existing approaches rely on document-level monolingual corpora in many different languages, sentence-level bilingual corpora,\footnote{In this p…

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationDocument Level Machine TranslationMachine Translation+2

The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining

2026-01-01 · Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani 외 arxiv

Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely believed to enable these abilities, details o…

KC4MT: A High-Quality Corpus for Multilingual Machine Translation

2022-06-01 · LREC 2022 6 · Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen 외

The multilingual parallel corpus is an important resource for many applications of natural language processing (NLP). For machine translation, the size and quality of the training corpus mainly affects the quality of the…

Machine TranslationSentenceTranslationVocal Bursts Intensity Prediction

DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational Recommendation

2021-09-18 · EMNLP 2021 11 · Zeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu 외

In this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendat…

Conversational Recommendation

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

2025-05-31 · Shaoxiong Ji, Zihao Li, Jaakko Paavola, Indraneil Paul 외

This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for m…

Translation