paper-with-me

Papers

Document Retrieval for Large Scale Content Analysis using Contextualized Dictionaries

2017-07-11 · Wiedemann Gregor, Niekler Andreas

This paper presents a procedure to retrieve subsets of relevant documents from large text collections for Content Analysis, e.g. in social sciences. Document retrieval for this purpose needs to take account of the fact that analysts often cannot describe their research objective with a small set of key terms, especially when dealing with theoretical or rather abstract research interests. Instead, it is much easier to define a set of paradigmatic documents which reflect topics of interest as well as targeted manner of speech. Thus, in contrast to classic information retrieval tasks we employ manually compiled collections of reference documents to compose large queries of several hundred key terms, called dictionaries. We extract dictionaries via Topic Models and also use co-occurrence data from reference collections. Evaluations show that the procedure improves retrieval results for this purpose compared to alternative methods of key term extraction as well as neglecting co-occurrence data.

📄 PDF Abstract BibTeX arXiv:1707.03217

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrievalTerm ExtractionTopic Models

Similar Papers 제목 키워드 기반

Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems

2026-01-08 · Uday Allu, Sonu Kedia, Tanmay Odapally, Biddwan Ahmed arxiv

Retrieval-Augmented Generation (RAG) systems critically depend on effective document chunking strategies to balance retrieval quality, latency, and operational cost. Traditional chunking approaches, such as fixed-size, r…

Text Generation

A Comparative Analysis of Retrievability and PageRank Measures

2023-11-17 · Aman Sinha, Priyanshu Raj Mall, Dwaipayan Roy

The accessibility of documents within a collection holds a pivotal role in Information Retrieval, signifying the ease of locating specific content in a collection of documents. This accessibility can be achieved via two …

Information RetrievalRetrieval

Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval

2024-06-11 · Adrià Molina, Oriol Ramos Terrades, Josep Lladós

This paper introduces Fetch-A-Set (FAS), a comprehensive benchmark tailored for legislative historical document analysis systems, addressing the challenges of large-scale document retrieval in historical contexts. The be…

Image RetrievalImage to textOptical Character Recognition (OCR)Retrieval

Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

2026-05-29 · Gyu-Hwung Cho, Youngjune Lee, Kiyoon Jeong, Siyoung Lee 외 arxiv

As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing attention; yet it still lacks a deployable system that lexically indexes v…

The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation

2025-04-11 · Zheng Zhang, Ning li, Qi Liu, Rui Li 외

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant document from external knowledge sources. By referencing this external knowledge, RAG effectively reduces the generation o…

FairnessHallucinationRAGRetrieval+1