paper-with-me

Papers

First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI

2024-12-12 · Sourav Banerjee, Anush Mahajan, Ayushi Agarwal, Eishkaran Singh

Natural Language Inference (NLI) tasks require identifying the relationship between sentence pairs, typically classified as entailment, contradiction, or neutrality. While the current state-of-the-art (SOTA) model, Entailment Few-Shot Learning (EFL), achieves a 93.1% accuracy on the Stanford Natural Language Inference (SNLI) dataset, further advancements are constrained by the dataset's limitations. To address this, we propose a novel approach leveraging synthetic data augmentation to enhance dataset diversity and complexity. We present UnitedSynT5, an advanced extension of EFL that leverages a T5-based generator to synthesize additional premise-hypothesis pairs, which are rigorously cleaned and integrated into the training data. These augmented examples are processed within the EFL framework, embedding labels directly into hypotheses for consistency. We train a GTR-T5-XL model on this expanded dataset, achieving a new benchmark of 94.7% accuracy on the SNLI dataset, 94.0% accuracy on the E-SNLI dataset, and 92.6% accuracy on the MultiNLI dataset, surpassing the previous SOTA models. This research demonstrates the potential of synthetic data augmentation in improving NLI models, offering a path forward for further advancements in natural language understanding tasks.

📄 PDF Abstract BibTeX arXiv:2412.09263

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDiversityFew-Shot LearningFew-Shot NLINatural Language InferenceNatural Language UnderstandingSentence

Similar Papers 제목 키워드 기반

DALL-E for Detection: Language-driven Compositional Image Synthesis for Object Detection

2022-06-20 · Yunhao Ge, Jiashu Xu, Brian Nlong Zhao, Neel Joshi 외

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-toimage synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach decouples training…

Image CaptioningImage GenerationObjectobject-detection+2

Factor-Conditioned Speaking-Style Captioning

2024-06-27 · Atsushi Ando, Takafumi Moriya, Shota Horiguchi, Ryo Masumura

This paper presents a novel speaking-style captioning method that generates diverse descriptions while accurately predicting speaking-style information. Conventional learning criteria directly use original captions that …

Diversity

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

2023-05-30 · Xingyu Fu, Sheng Zhang, Gukyeong Kwon, Pramuditha Perera 외

The open-ended Visual Question Answering (VQA) task requires AI models to jointly reason over visual and natural language inputs using world knowledge. Recently, pre-trained Language Models (PLM) such as GPT-3 have been …

Answer SelectionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)+1

AmbiPun: Generating Humorous Puns with Ambiguous Context

2022-05-04 · NAACL 2022 7 · Anirudh Mittal, Yufei Tian, Nanyun Peng

In this paper, we propose a simple yet effective way to generate pun sentences that does not require any training on existing puns. Our approach is inspired by humor theories that ambiguity comes from the context rather …

Reverse Dictionary

$\textsc{AmbiPun}$: Generating Humorous Puns with Ambiguous Context

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In this paper, we propose a simple yet effective way to generate pun sentences that does not require any training on existing puns. Our approach is inspired by humor theories that ambiguity comes from the context rather …

Reverse Dictionary