paper-with-me

홈 › Papers

Frustratingly Easy Natural Question Answering

2019-09-11 · Lin Pan, Rishav Chakravarti, Anthony Ferritto, Michael Glass, Alfio Gliozzo, Salim Roukos, Radu Florian, Avirup Sil

Existing literature on Question Answering (QA) mostly focuses on algorithmic novelty, data augmentation, or increasingly large pre-trained language models like XLNet and RoBERTa. Additionally, a lot of systems on the QA leaderboards do not have associated research documentation in order to successfully replicate their experiments. In this paper, we outline these algorithmic components such as Attention-over-Attention, coupled with data augmentation and ensembling strategies that have shown to yield state-of-the-art results on benchmark datasets like SQuAD, even achieving super-human performance. Contrary to these prior results, when we evaluate on the recently proposed Natural Questions benchmark dataset, we find that an incredibly simple approach of transfer learning from BERT outperforms the previous state-of-the-art system trained on 4 million more examples than ours by 1.9 F1 points. Adding ensembling strategies further improves that number by 2.3 F1 points.

📄 PDF Abstract BibTeX arXiv:1909.05286

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationNatural QuestionsQuestion AnsweringTransfer Learning

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
RoBERTa 설명 없음
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Weight Decay 설명 없음
SentencePiece 설명 없음

Similar Papers 제목 키워드 기반

Frustratingly Easy Uncertainty Estimation for Distribution Shift

2021-06-07 · Tiago Salvador, Vikram Voleti, Alexander Iannantuono, Adam Oberman

Distribution shift is an important concern in deep image classification, produced either by corruption of the source images, or a complete change, with the solution involving domain adaptation. While the primary goal is …

Domain Adaptationimage-classificationImage ClassificationUnsupervised Domain Adaptation

Frustratingly Easy Cross-Lingual Transfer for Transition-Based Dependency Parsing

2016-06-01 · NAACL 2016 6 · Oph{\'e}lie Lacroix, Lauriane Aufrant, Guillaume Wisniewski, Fran{\c{c}}ois Yvon
Cross-Lingual TransferDependency ParsingTransition-Based Dependency Parsing

Frustratingly Easy Label Projection for Cross-lingual Transfer

2022-11-28 · Yang Chen, Chao Jiang, Alan Ritter, Wei Xu

Translating training data into many languages has emerged as a practical solution for improving cross-lingual transfer. For tasks that involve span-level annotations, such as information extraction or question answering,…

Cross-Lingual NERCross-Lingual TransferEvent ExtractionNER+4

A Frustratingly Easy Improvement for Position Embeddings via Random Padding

2023-05-08 · Mingxu Tao, Yansong Feng, Dongyan Zhao

Position embeddings, encoding the positional relationships among tokens in text sequences, make great contributions to modeling local context features in Transformer-based pre-trained language models. However, in Extract…

Extractive Question-AnsweringPositionQuestion Answering

Answering questions by learning to rank - Learning to rank by answering questions

2019-11-01 · IJCNLP 2019 11 · George Sebastian Pirtoaca, Traian Rebedea, Stefan Ruseti

Answering multiple-choice questions in a setting in which no supporting documents are explicitly provided continues to stand as a core problem in natural language processing. The contribution of this article is two-fold.…

ARCLearning-To-RankMultiple-choice