paper-with-me

홈 › Papers

SPIRE: Structure-Preserving Interpretable Retrieval of Evidence

2026-02-12 · Mike Rainey, Umut Acar, Muhammed Sezer arxiv

Retrieval-augmented generation over semi-structured sources such as HTML is constrained by a mismatch between document structure and the flat, sequence-based interfaces of today's embedding and generative models. Retrieval pipelines often linearize documents into fixed-size chunks before indexing, which obscures section structure, lists, and tables, and makes it difficult to return small, citation-ready evidence without losing the surrounding context that makes it interpretable. We present a structure-aware retrieval pipeline that operates over tree-structured documents. The core idea is to represent candidates as subdocuments: precise, addressable selections that preserve structural identity while deferring the choice of surrounding context. We define a small set of document primitives--paths and path sets, subdocument extraction by pruning, and two contextualization mechanisms. Global contextualization adds the non-local scaffolding needed to make a selection intelligible (e.g., titles, headers, list and table structure). Local contextualization expands a seed selection within its structural neighborhood to obtain a compact, context-rich view under a target budget. Building on these primitives, we describe an embedding-based candidate generator that indexes sentence-seeded subdocuments and a query-time, document-aware aggregation step that amortizes shared structural context. We then introduce a contextual filtering stage that re-scores retrieved candidates using locally contextualized views. Across experiments on HTML question-answering benchmarks, we find that preserving structure while contextualizing selections yields higher-quality, more diverse citations under fixed budgets than strong passage-based baselines, while maintaining scalability.

📄 PDF Abstract BibTeX arXiv:2604.20849

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HyBIRD: Hyperbolic Bridge Retrieval and Diagnosis for Methodology Inspiration Retrieval

2026-05-29 · Yang Yang, Boyun Xu, Hao Fu, Jindong Li 외 arxiv

Methodology Inspiration Retrieval (MIR) asks a system to retrieve prior papers whose methods can inspire a new research proposal. Unlike general scientific retrieval, the central challenge is not topical similarity but w…

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

2026-05-29 · Yating Pan, Jiajun Zhang, Jun Wang, Qi Su arxiv

LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantitative signals. Humanities scholarship, however, requires interpretiv…

Towards Agentic Defect Reasoning: A Graph-Assisted Retrieval Framework for Laser Powder Bed Fusion

2026-04-05 · Muhammad Rizwan Awan, Volker Pickert, Muhammad Waqar Ashraf, Saleh Ali 외 arxiv

Laser Powder Bed Fusion (LPBF) is highly sensitive to process parameters, which influence defect formation through complex thermal and fluid mechanisms. However, defect-related knowledge is dispersed across the literatur…

N2N-GQA: Noise-to-Narrative for Graph-Based Table-Text Question Answering Using LLMs

2026-01-10 · Mohamed Sharafath, Aravindh Annamalai, Ganesh Murugan, Aravindakumar Venugopalan arxiv

Multi-hop question answering over hybrid table-text data requires retrieving and reasoning across multiple evidence pieces from large corpora, but standard Retrieval-Augmented Generation (RAG) pipelines process documents…

Multi-hop Question Answering

Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG

2026-05-31 · Francielle Vargas, João Robiatti, Diego Alves, Lucas Pascotti Valem 외 arxiv

Ensuring factuality and interpretability in RAG remains an open and urgent problem. We introduce Contrastive Evidence Rationale Attention (CERA), the first retrieval framework to employ subjectivity-based hard negative s…

Contrastive Learning