paper-with-me

Papers

MCiteBench: A Multimodal Benchmark for Generating Text with Citations

2025-03-04 · Caiyu Hu, Yikai Zhang, Tinghui Zhu, Yiwei Ye, Yanghua Xiao

Multimodal Large Language Models (MLLMs) have advanced in integrating diverse modalities but frequently suffer from hallucination. A promising solution to mitigate this issue is to generate text with citations, providing a transparent chain for verification. However, existing work primarily focuses on generating citations for text-only content, leaving the challenges of multimodal scenarios largely unexplored. In this paper, we introduce MCiteBench, the first benchmark designed to assess the ability of MLLMs to generate text with citations in multimodal contexts. Our benchmark comprises data derived from academic papers and review-rebuttal interactions, featuring diverse information sources and multimodal content. Experimental results reveal that MLLMs struggle to ground their outputs reliably when handling multimodal input. Further analysis uncovers a systematic modality bias and reveals how models internally rely on different sources when generating citations, offering insights into model behavior and guiding future directions for multimodal citation tasks.

📄 PDF Abstract BibTeX arXiv:2503.02589

Code (1)

caiyuhu/MCiteBench 공식 구현

Tasks

HallucinationText Generation

Similar Papers 제목 키워드 기반

MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering

2025-11-15 · Seokwon Song, Minsu Park, Gunhee Kim arxiv

Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-onl…

Mitigating Contextual BiasVisual Question Answering

Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation

2025-09-25 · Guo Chen, Qiuyuan Li, Qiuxian Li, Hongliang Dai 외 arxiv

In retrieval-augmented generation (RAG) question answering systems, generating citations for large language model (LLM) outputs enhances verifiability and helps users identify potential hallucinations. However, we observ…

Question Answering

Learning Fine-Grained Grounded Citations for Attributed Large Language Models

2024-08-08 · Lei Huang, Xiaocheng Feng, Weitao Ma, Yuxuan Gu 외

Despite the impressive performance on information-seeking tasks, large language models (LLMs) still struggle with hallucinations. Attributed LLMs, which augment generated text with in-line citations, have shown potential…

In-Context Learning

Multimodal Fact-Level Attribution for Verifiable Reasoning

2026-02-12 · David Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin 외 arxiv

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sourc…

Multimodal Reasoning

SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models

2025-02-13 · Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, Zhaofeng Wu 외

We introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly…

Long Form Question AnsweringQuestion AnsweringSentence