Do Current Natural Language Inference Models Truly Understand Sentences? Insights from Simple Sentences
Natural language inference (NLI) is a task to infer the relationship between a premise and a hypothesis (e.g. entailment, neutral, or contradiction), and transformer-based models perform well on current NLI datasets such as MNLI and SNLI. Nevertheless, given the complexity of the task, especially the complexity of the sentences used for model evaluations, it remains controversial whether these models can truly infer the meaning of sentences or they simply guess the answer via non-humanlike heuristics. Here, we reduce the complexity of the task using two approaches. The first approach simplifies the relationship between the premise and hypothesis by making them unrelated. A test set, referred to as Random Pair, is constructed by randomly pairing premises and hypotheses in MNLI/SNLI. Models fine-tuned on MNLI/SNLI identify a large proportion (up to 77.6%) of these unrelated statements as being contradictory. Models fine-tuned on SICK, a dataset that included unrelated premise-hypothesis pairs, perform well on Random Pair. The second approach simplifies the task by constraining the premises/hypotheses to be syntactically/semantically simple sentences. A new test set, referred to as Simple Pair, is constructed using simple sentences, such as short SVO sentences, and basic conjunction sentences. We find that models fine-tuned on MNLI/SNLI generally fail to understand these simple sentences, but their performance can be boosted by re-fine-tuning the models using only a few hundreds of samples from SICK. All models tested here, however, fail to understand the fundamental compositional binding relation between a subject and a predicate (up to ~100% error rate) for basic conjunction sentences. Taken together, the results show that models achieving high accuracy on mainstream datasets can still lack basic sentence comprehension capacity, and datasets discouraging non-humanlike heuristics are required to build more robust NLI models.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language InferenceSentenceSimilar Papers 제목 키워드 기반
Can 3D Vision-Language Models Truly Understand Natural Language?
Rapid advancements in 3D vision-language (3D-VL) tasks have opened up new avenues for human interaction with embodied agents or robots using natural language. Despite this progress, we find a notable limitation: existing…
BenchmarkingDiversityDKPLM: Decomposable Knowledge-enhanced Pre-trained Language Model for Natural Language Understanding
Knowledge-Enhanced Pre-trained Language Models (KEPLMs) are pre-trained models with relation triples injecting from knowledge graphs to improve language understanding abilities. To guarantee effective knowledge injection…
Knowledge GraphsKnowledge ProbingLanguage ModelingLanguage Modelling+2LLMs' Understanding of Natural Language Revealed
Large language models (LLMs) are the result of a massive experiment in bottom-up, data-driven reverse engineering of language at scale. Despite their utility in a number of downstream NLP tasks, ample research has shown …
MemorizationText GenerationExploring the Role of Reasoning Structures for Constructing Proofs in Multi-Step Natural Language Reasoning with Large Language Models
When performing complex multi-step reasoning tasks, the ability of Large Language Models (LLMs) to derive structured intermediate proof steps is important for ensuring that the models truly perform the desired reasoning …
In-Context LearningKnowledge-driven Natural Language Understanding of English Text and its Applications
Understanding the meaning of a text is a fundamental challenge of natural language understanding (NLU) research. An ideal NLU system should process a language in a way that is not exclusive to a single task or a dataset.…
Natural Language UnderstandingQuestion Answering