paper-with-me

홈 › Papers

Developing a Dataset for Evaluating Approaches for Document Expansion with Images

2016-05-01 · LREC 2016 5 · Debasis Ganguly, Iacer Calixto, Gareth Jones

Motivated by the adage that a {``}picture is worth a thousand words{''} it can be reasoned that automatically enriching the textual content of a document with relevant images can increase the readability of a document. Moreover, features extracted from the additional image data inserted into the textual content of a document may, in principle, be also be used by a retrieval engine to better match the topic of a document with that of a given query. In this paper, we describe our approach of building a ground truth dataset to enable further research into automatic addition of relevant images to text documents. The dataset is comprised of the official ImageCLEF 2010 collection (a collection of images with textual metadata) to serve as the images available for automatic enrichment of text, a set of 25 benchmark documents that are to be enriched, which in this case are children{'}s short stories, and a set of manually judged relevant images for each query story obtained by the standard procedure of depth pooling. We use this benchmark dataset to evaluate the effectiveness of standard information retrieval methods as simple baselines for this task. The results indicate that using the whole story as a weighted query, where the weight of each query term is its tf-idf value, achieves an precision of 0:1714 within the top 5 retrieved images on an average.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

AdaQE-CG: Adaptive Query Expansion for Web-Scale Generative AI Model and Data Card Generation

2026-03-16 · Haoxuan Zhang, Ruochi Li, Zhenni Liang, Mehri Sattari 외 arxiv

Transparent and standardized documentation is essential for building trustworthy generative AI (GAI) systems. However, existing automated methods for generating model and data cards still face three major challenges: (i)…

Information Extraction

NovAScore: A New Automated Metric for Evaluating Document Level Novelty

2024-09-14 · Lin Ai, Ziwei Gong, Harshsaiprasad Deshpande, Alexander Johnson 외

The rapid expansion of online content has intensified the issue of information redundancy, underscoring the need for solutions that can identify genuinely new information. Despite this challenge, the research community h…

Novelty Detection

Experiments on Manual Thesaurus based Query Expansion for Ad-hoc Monolingual Gujarati Information Retrieval Tasks

2020-01-18 · Hardik Joshi, Jyoti Pareek

In this paper, we present the experimental work done on Query Expansion (QE) for retrieval tasks of Gujarati text documents. In information retrieval, it is very difficult to estimate the exact user need, query expansion…

Information RetrievalRetrieval

Semantic Evolutionary Concept Distances for Effective Information Retrieval in Query Expansion

2017-01-19 · Valentina Franzoni, Yuanxi Li, Clement H. C. Leung, Alfredo Milani

In this work several semantic approaches to concept-based query expansion and reranking schemes are studied and compared with different ontology-based expansion methods in web document search and retrieval. In particular…

Information RetrievalRerankingRetrieval

Query Expansion Strategy based on Pseudo Relevance Feedback and Term Weight Scheme for Monolingual Retrieval

2015-02-18 · Vaidyanathan Rekha, Das Sujoy, Srivastava Namita

Query Expansion using Pseudo Relevance Feedback is a useful and a popular technique for reformulating the query. In our proposed query expansion method, we assume that relevant information can be found within a document …

Retrieval