paper-with-me

Natural Language Inference

33개 벤치마크 · 논문 2,097편 · 이 태스크의 논문 보기 →

Benchmarks

SNLI

결과 98개

RTE

결과 90개

MultiNLI

결과 67개

QNLI

결과 43개

ANLI test

결과 25개

WNLI

결과 23개

LiDiRus

결과 22개

RCB

결과 22개

TERRa

결과 22개

CommitmentBank

결과 20개

SciTail

결과 13개

FarsTail

결과 10개

MultiNLI Dev

결과 10개

MedNLI

결과 7개

XNLI French

결과 6개

V-SNLI

결과 3개

XNLI Chinese

결과 3개

XNLI Chinese Dev

결과 3개

e-SNLI

결과 3개

JamPatoisNLI

결과 2개

AX

결과 1개

BioNLI

결과 1개

HANS

결과 1개

KUAKE-QQR

결과 1개

KUAKE-QTR

결과 1개

MED

결과 1개

MRPC

결과 1개

Probability words NLI

결과 1개

Quora Question Pairs

결과 1개

SICK

결과 1개

TabFact

결과 1개

XWINO

결과 1개

Most implemented

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

Papers

Cascaded Batch Prompting

2026-08-27 · Sho Hoshino, Peinan Zhang arxiv

Although batch prompting makes large language model inference more efficient by processing multiple instances simultaneously, it suffers from unpredictable downstream task performance. We propose cascaded batch prompting…

Natural Language InferenceQuestion Answering

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

2026-08-21 · Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem 외 arxiv

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance do…

Natural Language Inference

Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI

2026-08-19 · Toheeb Ogunade arxiv

We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural langua…

Natural Language Inference

Consensus Measures for Unstructured Biomedical Text Annotations

2026-08-04 · Pascal Wullschleger, Christian Kreis, Martin A. Walter, Marc Pouly 외 arxiv

Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard …

Natural Language Inference

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

2026-07-22 · Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean 외 arxiv

AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free fram…

Natural Language InferenceMathematical ReasoningQuestion Answering

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

2026-07-17 · Haram Choi arxiv

Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SN…

Natural Language Inference

전체 2,097편 보기 →