paper-with-me

Papers

Graph Evidence Is Not Enough: Diagnosing Native Decoder Use in Graph-Augmented LLMs

2026-08-31 · Xiaoyu Guo, Pengcheng Chen, Jiong Yu, Yi Lu, Yaohua Wang, Ziyang Li arxiv

Graph-augmented large language models often assume that graph evidence produced by external computation and placed in the input can be used by the native decoder. We test this assumption with HopQA, a deliberately bounded diagnostic that asks for the shortest-hop distance between two query nodes. Because the answer is a small integer and the target is purely topological, failure cannot be dismissed as open-ended generation or ambiguous evaluation. Yet existing graph-augmented baselines still fail on this setting, showing that providing graph evidence is not the same as making it usable. We introduce an intervention triangle with three matched conditions: readable graph evidence, shuffled graph evidence, and no-graph input. This separates evidence inclusion, structural readability, and decoder-usable topology. Guided by this diagnosis, we present S$^2$GE as an instance showing that diagnosis-driven interface design can improve native decoder usability. S$^2$GE uses query-aware sampling, endpoint and proximity-based ordering, and structure-preserving alignment. Across DBLP, Biomedical, GoodReads, and PubMed, S$^2$GE achieves strict exact-match scores of $36.5\%$, $57.8\%$, $76.6\%$, and $52.0\%$, improving over the strongest native-generation baseline by $53.5$ points on average. The interventions further reveal harmful-shuffle, shuffle-robust, and no-graph-saturated regimes.

📄 PDF Abstract BibTeX arXiv:2608.30437

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis

2026-01-10 · Wenting Chen, Zhongrui Zhu, Guolin Huang, Wenxuan Wang arxiv

Despite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis--relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical c…

Causal Inference

Lite Unified Modeling for Discriminative Reading Comprehension

2022-03-26 · ACL 2022 5 · Yilin Zhao, Hai Zhao, Libin Shen, Yinggong Zhao

As a broad and major category in machine reading comprehension (MRC), the generalized goal of discriminative MRC is answer prediction from the given materials. However, the focuses of various discriminative MRC tasks may…

DecoderMachine Reading ComprehensionMulti-Choice MRCPOS+1

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification

2026-05-26 · Jingxi Qiu, Zeyu Han, Cheng Huang arxiv

Evidence absence is not evidence insufficiency, but fact verification benchmarks can make them observationally similar. The Not Enough Information (NEI) label is often operationalized through different evidence condition…

Fact Verification

SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature

2026-01-15 · Yiming Ren, Junjie Wang, Yuxin Meng, Yihang Shi 외 arxiv

Evaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic "Needle-In-A-Haystack" tests often reward answer matching without re…

COVID-CT-Dataset: A CT Scan Dataset about COVID-19

2020-03-30 · Xingyi Yang, Xuehai He, Jinyu Zhao, Yichen Zhang 외

During the outbreak time of COVID-19, computed tomography (CT) is a useful manner for diagnosing COVID-19 patients. Due to privacy issues, publicly available COVID-19 CT datasets are highly difficult to obtain, which hin…

Computed Tomography (CT)COVID-19 DiagnosisMulti-Task LearningSelf-Supervised Learning