paper-with-me

홈 › Papers

LINDA: Unsupervised Learning to Interpolate in Natural Language Processing

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Despite the success of mixup in data augmentation, its applicability to natural language processing (NLP) tasks has been limited due to the discrete and variable-length nature of natural languages. Recent studies have thus relied on domain-specific heuristics and manually crafted resources, such as dictionaries, in order to apply mixup in NLP. In this paper, we instead propose an unsupervised learning approach to text interpolation for the purpose of data augmentation, to which we refer as 'Learning to INterpolate for Data Augmentation' (LINDA), that does not require any heuristics nor manually crafted resources but learns to interpolate between any pair of natural language sentences over a natural language manifold. After empirically demonstrating the LINDA's interpolation capability, we show that LINDA indeed allows us to seamlessly apply mixup in NLP and leads to better generalization in text classification both in-domain and out-of-domain.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentationtext-classificationText Classification

Methods 이 논문이 사용한 방법론

Mixup Mixup is a data augmentation technique that generates a weighted combination of random image pairs from the training data. Given two images and their ground truth labels:…

Similar Papers 제목 키워드 기반

LINDA: Unsupervised Learning to Interpolate in Natural Language Processing

2021-12-28 · Yekyung Kim, Seohyeong Jeong, Kyunghyun Cho

Despite the success of mixup in data augmentation, its applicability to natural language processing (NLP) tasks has been limited due to the discrete and variable-length nature of natural languages. Recent studies have th…

Data Augmentationtext-classificationText Classification

What's the Problem, Linda? The Conjunction Fallacy as a Fairness Problem

2023-05-16 · Jose Alvarez Colmenares

The field of Artificial Intelligence (AI) is focusing on creating automated decision-making (ADM) systems that operate as close as possible to human-like intelligence. This effort has pushed AI researchers into exploring…

Decision MakingFairness

Speech Technology Services for Oral History Research

2024-04-26 · Christoph Draxler, Henk van den Heuvel, Arjan van Hessen, Pavel Ircing 외

Oral history is about oral sources of witnesses and commentors on historical events. Speech technology is an important instrument to process such recordings in order to obtain transcription and further enhancements to st…

AP-OOD: Attention Pooling for Out-of-Distribution Detection

2026-02-05 · Claus Hofmann, Christian Huber, Bernhard Lehner, Daniel Klotz 외 arxiv

Out-of-distribution (OOD) detection, which maps high-dimensional data into a scalar OOD score, is critical for the reliable deployment of machine learning models. A key challenge in recent research is how to effectively …

Out-of-Distribution Detection

MELINDA: A Multimodal Dataset for Biomedical Experiment Method Classification

2020-12-16 · Te-Lin Wu, Shikhar Singh, Sayan Paul, Gully Burns 외

We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing cu…

ClassificationGeneral Classification