paper-with-me

Papers Natural Language Inference

“Natural Language Inference” 태그가 달린 논문 2,097편 · 필터 해제

Cascaded Batch Prompting

2026-08-27 · Sho Hoshino, Peinan Zhang arxiv

Although batch prompting makes large language model inference more efficient by processing multiple instances simultaneously, it suffers from unpredictable downstream task performance. We propose cascaded batch prompting…

Natural Language InferenceQuestion Answering

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

2026-08-21 · Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem 외 arxiv

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance do…

Natural Language Inference

Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI

2026-08-19 · Toheeb Ogunade arxiv

We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural langua…

Natural Language Inference

Consensus Measures for Unstructured Biomedical Text Annotations

2026-08-04 · Pascal Wullschleger, Christian Kreis, Martin A. Walter, Marc Pouly 외 arxiv

Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard …

Natural Language Inference

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

2026-07-22 · Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean 외 arxiv

AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free fram…

Natural Language InferenceMathematical ReasoningQuestion Answering

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

2026-07-17 · Haram Choi arxiv

Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SN…

Natural Language Inference

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

2026-07-14 · Hielke Muizelaar, Giulia Rivetti, Marco Spruit, Marcel Haas arxiv

BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains …

Natural Language InferencePart-Of-Speech TaggingHate Speech DetectionQuestion Answering

Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP)

2026-07-05 · Mohamed Aly Bouke arxiv

Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why. To close this gap…

Natural Language InferenceQuestion Answering

Where do LLMs Fall Short in CBT-Guided Affective Reasoning?

2026-07-03 · Vaishnavi Sinha, Pooja Guttal, Pranay Deep Reddy Katike, Vishal Sinha 외 arxiv

Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction between cognitive and behavioral factors. However, out-of-the-box LLMs respond fluen…

Natural Language Inference

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

2026-07-02 · Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah, Shahriar Noroozizadeh arxiv

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit …

Natural Language InferenceSemantic Similarity

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

2026-06-29 · Kriti Faujdar, Smit Kadvani arxiv

Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-intensive inference, proprietary API calls, or white-box access to the g…

Natural Language InferenceSemantic SimilarityQuestion Answering

MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation

2026-06-20 · Mohammed Rawhani, Dervis Karaboga, Ozkan Ufuk Nalbantoglu, Alper Basturk 외 arxiv

Pre-trained language models struggle when applied to new domains, as full fine-tuning is computationally expensive and prone to catastrophic forgetting. This study addresses this challenge by presenting a novel parameter…

Unsupervised Domain AdaptationNatural Language Inference

Evaluating LLM Personalization via Semantic Constraint Verification

2026-06-15 · Xuran Li, Guanqin Zhang, Imran Razzak, Hakim Hacid 외 arxiv

Current evaluation paradigms for Large Language Model (LLM) personalization rely heavily on brittle surface-matching metrics or computationally expensive LLM-as-a-judge protocols, both of which lack interpretability. To …

Natural Language Inference

Can News Predict the Market? Limits of Zero-Shot Financial NLP and the Role of Explainable AI

2026-06-10 · Ali M Karaoglu, Shreyank N Gowda arxiv

Can financial news reliably predict short-term stock movements? Despite advances in large language models, this question remains unresolved. We revisit this problem using a zero-shot natural language processing framework…

Natural Language Inference

TruthSplit: Operationalizing Conditional Validity in Arguments Through Multi-Perspective Reasoning

2026-06-08 · Benjamin Stieger, Maximilian Terberger, Thomas Huber, Christina Niklaus arxiv

We present TruthSplit, an interactive system for multi-perspective argument analysis. Existing argumentation tools typically analyze properties of the argument itself, such as structure, quality, stance, or persuasivenes…

Natural Language Inference

SHALA-LLM: Smartly Handling Ambiguous Labels in Aligning LLMs

2026-06-03 · Jingyao Wu, Ashley Wang, Keane Ong, Paul Pu Liang 외 arxiv

Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challenging disagreements across human annotator…

Natural Language InferenceReinforcement LearningEmotion Recognition

Sample-Size Scaling of the African Languages NLI Evaluation

2026-06-02 · Anuj Tiwari, Oluwapelumi Ogunremu, Terry Oko-odion, Jesujuwon Egbewale 외 arxiv

African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance. The study is a systematic sample-size scaling study of natural l…

Natural Language Inference

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

2026-06-02 · Peerawat Chomphooyod, Jian Gang Ngui, Yosephine Susanto, Attapol T. Rutherford 외 arxiv

Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NLI benchmarks are largely Western-centric, translation-derived, or mo…

Natural Language Inference

From Script to Semantics: Prompting Strategies for African NLI

2026-06-02 · Anuj Tiwari, Terry Oko-odion, Hannah Nwokocha arxiv

Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning. …

Natural Language Inference

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

2026-06-01 · Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani 외 arxiv

Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datas…

Natural Language Inference
1–20 / 2,097 다음 →