paper-with-me

Papers

How Well Do Large Language Models Truly Ground?

2023-11-15 · Hyunji Lee, Sejune Joo, Chaeeun Kim, Joel Jang, Doyoung Kim, Kyoung-Woon On, Minjoon Seo

To reduce issues like hallucinations and lack of control in Large Language Models (LLMs), a common method is to generate responses by grounding on external contexts given as input, known as knowledge-augmented models. However, previous research often narrowly defines "grounding" as just having the correct answer, which does not ensure the reliability of the entire response. To overcome this, we propose a stricter definition of grounding: a model is truly grounded if it (1) fully utilizes the necessary knowledge from the provided context, and (2) stays within the limits of that knowledge. We introduce a new dataset and a grounding metric to evaluate model capability under the definition. We perform experiments across 25 LLMs of different sizes and training methods and provide insights into factors that influence grounding performance. Our findings contribute to a better understanding of how to improve grounding capabilities and suggest an area of improvement toward more reliable and controllable LLM applications.

📄 PDF Abstract BibTeX arXiv:2311.09069

Code (1)

kaistai/how-well-do-llms-truly-ground 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Conceptual Grounding Constraints for Truly Robust Biomedical Name Representations

2021-04-01 · EACL 2021 2 · Pieter Fivez, Simon Suster, Walter Daelemans

Effective representation of biomedical names for downstream NLP tasks requires the encoding of both lexical as well as domain-specific semantic information. Ideally, the synonymy and semantic relatedness of names should …

Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages

2020-04-28 · Katharina Kann, Ophélie Lacroix, Anders Søgaard

Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision - e.g., cross-lingual transfer, type-level supervision, or a combination thereof - have been report…

Cross-Lingual TransferPOSPOS Tagging

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

2025-11-24 · Federico Felizzi, Olivia Riccomi, Michele Ferramola, Francesco Andrea Causio 외 arxiv

Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs dem…

Visual Question AnsweringVisual Grounding

Metaphors We Compute By: A Computational Audit of Cultural Translation vs. Thinking in LLMs

2026-04-06 · Yuan Chang, Jiaming Qu, Zhu Li arxiv

Large language models (LLMs) are often described as multilingual because they can understand and respond in many languages. However, speaking a language is not the same as reasoning within a culture. This distinction mot…

SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation

2025-12-01 · Sheng Liu, Tianyu Luan, Phani Nuney, Xuelu Feng 외 arxiv

3D generation and reconstruction techniques have been widely used in computer games, film, and other content creation areas. As the application grows, there is a growing demand for 3D shapes that look truly realistic. Tr…

3D Generation