paper-with-me

홈 › Papers

A Benchmark Suite for Template Detection and Content Extraction

2014-09-22 · Julián Alarte, Josep Silva

Template detection and content extraction are two of the main areas of information retrieval applied to the Web. They perform different analyses over the structure and content of webpages to extract some part of the document. However, their objective is different. While template detection identifies the template of a webpage (usually comparing with other webpages of the same website), content extraction identifies the main content of the webpage discarding the other part. Therefore, they are somehow complementary, because the main content is not part of the template. It has been measured that templates represent between 40% and 50% of data on the Web. Therefore, identifying templates is essential for indexing tasks because templates usually contain irrelevant information such as advertisements, menus and banners. Processing and storing this information is likely to lead to a waste of resources (storage space, bandwidth, etc.). Similarly, identifying the main content is essential for many information retrieval tasks. In this paper, we present a benchmark suite to test different approaches for template detection and content extraction. The suite is public, and it contains real heterogeneous webpages that have been labelled so that different techniques can be suitable (and automatically) compared.

📄 PDF Abstract BibTeX arXiv:1409.6182

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

The Impact of Main Content Extraction on Near-Duplicate Detection

2021-11-21 · Maik Fröbe, Matthias Hagen, Janek Bevendorff, Michael Völske 외

Commercial web search engines employ near-duplicate detection to ensure that users see each relevant result only once, albeit the underlying web crawls typically include (near-)duplicates of many web pages. We revisit th…

Information RetrievalRetrieval

Extraction of Templates from Phrases Using Sequence Binary Decision Diagrams

2020-01-28 · Daiki Hirano, Kumiko Tanaka-Ishii, Andrew Finch

The extraction of templates such as ``regard X as Y'' from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-f…

MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling

2024-06-18 · Philipp Seeberger, Dominik Wagner, Korbinian Riedhammer

With the advancement of multimedia technologies, news documents and user-generated content are often represented as multiple modalities, making Multimedia Event Extraction (MEE) an increasingly important challenge. Howev…

Data AugmentationEvent Argument ExtractionEvent Extraction

Similar Document Template Matching Algorithm

2023-11-21 · Harshitha Yenigalla, Bommareddy Revanth Srinivasa Reddy, Batta Venkata Rahul, Nannapuraju Hemanth Raju

This study outlines a comprehensive methodology for verifying medical documents, integrating advanced techniques in template extraction, comparison, and fraud detection. It begins with template extraction using sophistic…

Fraud DetectionOptical Character Recognition (OCR)SSIMTemplate Matching

AI Sound Recognition on Asthma Medication Adherence: Evaluation With the RDA Benchmark Suite

2023-02-08 · IEEE Access 2023 2 · Dimitris Nikos Fakotakis, Stavros Nousias, Gerasimos Arvanitis, Evangelia I. Zacharaki 외

Asthma is a common, usually long-term respiratory disease with negative impact on global society and economy. Treatment involves using medical devices (inhalers) that distribute medication to the airways and its efficien…

BenchmarkingManagement