paper-with-me

Papers

CODEC: Complex Document and Entity Collection

2022-05-09 · Iain Mackie, Paul Owoicho, Carlos Gemmell, Sophie Fischer, Sean MacAvaney, Jeffrey Dalton

CODEC is a document and entity ranking benchmark that focuses on complex research topics. We target essay-style information needs of social science researchers, i.e. "How has the UK's Open Banking Regulation benefited Challenger Banks?". CODEC includes 42 topics developed by researchers and a new focused web corpus with semantic annotations including entity links. This resource includes expert judgments on 17,509 documents and entities (416.9 per topic) from diverse automatic and interactive manual runs. The manual runs include 387 query reformulations, providing data for query performance prediction and automatic rewriting evaluation. CODEC includes analysis of state-of-the-art systems, including dense retrieval and neural re-ranking. The results show the topics are challenging with headroom for document and entity ranking improvement. Query expansion with entity information shows significant gains in document ranking, demonstrating the resource's value for evaluating and improving entity-oriented search. We also show that the manual query reformulations significantly improve document ranking and entity ranking performance. Overall, CODEC provides challenging research topics to support the development and evaluation of entity-centric search methods.

📄 PDF Abstract BibTeX arXiv:2205.04546

Code (2)

grill-lab/codec 공식 구현
mihirs16/multi-stage-retrieval-using-rm3-and-t5

Tasks

Document RankingRe-RankingRetrieval

Similar Papers 제목 키워드 기반

Query-Specific Knowledge Graphs for Complex Finance Topics

2022-11-08 · Iain Mackie, Jeffrey Dalton

Across the financial domain, researchers answer complex questions by extensively "searching" for relevant information to generate long-form reports. This workshop paper discusses automating the construction of query-spec…

Document RankingKnowledge GraphsRetrieval

Toward Sub-1 kB Identity-Preserving Face Compression: A Benchmark of Codecs, a Custom Learned Codec, and Studies of Resolution, Demographic Fairness, Recompression, and Adversarial Robustness

2026-08-24 · Petr Hurtik, Jakub Sochor arxiv

Storing face images under a hard sub-kilobyte budget, as required for identity documents, smart-card biometrics and bandwidth-constrained verification, forces a codec to discard most of the signal while keeping what a fa…

Adversarial Robustness

MMEAD: MS MARCO Entity Annotations and Disambiguations

2023-09-14 · Chris Kamphuis, Aileen Lin, Siwen Yang, Jimmy Lin 외

MMEAD, or MS MARCO Entity Annotations and Disambiguations, is a resource for entity links for the MS MARCO datasets. We specify a format to store and share links for both document and passage collections of MS MARCO. Fol…

Entity Embeddings

Entity Labels Are Not Entity Signals: A Framework for Observable Relevance in Document Re-Ranking

2026-06-14 · Utshab Kumar Ghosh, Shubham Chatterjee arxiv

Entity-aware document retrieval uses query-associated entities as ranking signals, assuming that semantically relevant entities are also useful retrieval signals. We show this assumption is insufficient- and explain why.…

Adaptive Latent Entity Expansion for Document Retrieval

2023-06-29 · Iain Mackie, Shubham Chatterjee, Sean MacAvaney, Jeffrey Dalton

Despite considerable progress in neural relevance ranking techniques, search engines still struggle to process complex queries effectively - both in terms of precision and recall. Sparse and dense Pseudo-Relevance Feedba…

Re-RankingRetrieval