paper-with-me

홈 › Papers

From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines

2026-04-15 · Sunkyung Lee, Jihye Back, Donghyeon Jeon, Soonhwan Kwon, Moonkwon Kim, Inho Kang, Jongwuk Lee arxiv

Generative information retrieval (GenIR) formulates the retrieval process as a text-to-text generation task, leveraging the vast knowledge of large language models. However, existing works primarily optimize for relevance while often overlooking document trustworthiness. This is critical in high-stakes domains like healthcare and finance, where relying solely on semantic relevance risks retrieving unreliable information. To address this, we propose an Authority-aware Generative Retriever (AuthGR), the first framework that incorporates authority into GenIR. AuthGR consists of three key components: (i) Multimodal Authority Scoring, which employs a vision-language model to quantify authority from textual and visual cues; (ii) a Three-stage Training Pipeline to progressively instill authority awareness into the retriever; and (iii) a Hybrid Ensemble Pipeline for robust deployment. Offline evaluations demonstrate that AuthGR successfully enhances both authority and accuracy, with our 3B model matching a 14B baseline. Crucially, large-scale online A/B tests and human evaluations conducted on the commercial web search platform confirm significant improvements in real-world user engagement and reliability.

📄 PDF Abstract BibTeX arXiv:2604.13468

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalText Generation

Similar Papers 제목 키워드 기반

Controlling Authority Retrieval: A Missing Retrieval Objective for Authority-Governed Knowledge

2026-04-15 · Andre Bacellar arxiv

In law, regulatory regimes for pharmaceuticals and software security, newer authorities can revoke older established ones even when semantically distant. We call this CAR: retrieving the currently active authority fronti…

Joint Modeling of Topics, Citations, and Topical Authority in Academic Corpora

2017-06-02 · TACL 2017 1 · Jooyeon Kim, Dongwoo Kim, Alice Oh

Much of scientific progress stems from previously published findings, but searching through the vast sea of scientific publications is difficult. We often rely on metrics of scholarly authority to find the prominent auth…

An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?

2026-03-11 · Jennifer D'Souza, Sameer Sadruddin, Maximilian Kähler, Andrea Salfinger 외 arxiv

Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with the Integrated Authority File (GND), plus …

Multi-Label Text ClassificationMulti-Label Classification

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

2026-08-17 · Junjie Chu, Ye Leng, Mingjie Li, Yun Shen 외 arxiv

Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically optimized pages visibility disproportionate to th…

Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority

2026-07-23 · Shen Xu arxiv

AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence selection, and knowledge-intensive reasoning. Yet importance is often reduced to a single score derived from eithe…

Entity Alignment