Natural Language Inference
33개 벤치마크 · 논문 2,097편 · 이 태스크의 논문 보기 →
Benchmarks
SNLI
RTE
MultiNLI
QNLI
ANLI test
WNLI
LiDiRus
RCB
TERRa
CommitmentBank
SciTail
FarsTail
MultiNLI Dev
MedNLI
XNLI French
V-SNLI
XNLI Chinese
XNLI Chinese Dev
e-SNLI
JamPatoisNLI
AX
BioNLI
HANS
KUAKE-QQR
KUAKE-QTR
MED
MRPC
Probability words NLI
Quora Question Pairs
SICK
TabFact
XWINO
Most implemented
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
A Structured Self-attentive Sentence Embedding
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Papers
Cascaded Batch Prompting
Although batch prompting makes large language model inference more efficient by processing multiple instances simultaneously, it suffers from unpredictable downstream task performance. We propose cascaded batch prompting…
Natural Language InferenceQuestion AnsweringTrustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance do…
Natural Language InferenceStructure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI
We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural langua…
Natural Language InferenceConsensus Measures for Unstructured Biomedical Text Annotations
Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard …
Natural Language InferenceReference-Free Evaluation of Reasoning in Open-Ended Question Answering
AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free fram…
Natural Language InferenceMathematical ReasoningQuestion AnsweringHow Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI
Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SN…
Natural Language Inference