paper-with-me

Papers

Multi-Record Web Page Information Extraction From News Websites

2025-02-20 · Alexander Kustenkov, Maksim Varlamov, Alexander Yatskov

In this paper, we focused on the problem of extracting information from web pages containing many records, a task of growing importance in the era of massive web data. Recently, the development of neural network methods has improved the quality of information extraction from web pages. Nevertheless, most of the research and datasets are aimed at studying detailed pages. This has left multi-record "list pages" relatively understudied, despite their widespread presence and practical significance. To address this gap, we created a large-scale, open-access dataset specifically designed for list pages. This is the first dataset for this task in the Russian language. Our dataset contains 13,120 web pages with news lists, significantly exceeding existing datasets in both scale and complexity. Our dataset contains attributes of various types, including optional and multi-valued, providing a realistic representation of real-world list pages. These features make our dataset a valuable resource for studying information extraction from pages containing many records. Furthermore, we proposed our own multi-stage information extraction methods. In this work, we explore and demonstrate several strategies for applying MarkupLM to the specific challenges of multi-record web pages. Our experiments validate the advantages of our methods. By releasing our dataset to the public, we aim to advance the field of information extraction from multi-record pages.

📄 PDF Abstract BibTeX arXiv:2502.14625

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multilingual Attribute Extraction from News Web Pages

2025-02-04 · Pavel Bedrin, Maksim Varlamov, Alexander Yatskov

This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high efficacy in extracting information from semi-s…

AttributeAttribute Extraction

End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940

2024-04-30 · Thomas Constum, Lucas Preel, Théo Larcher, Pierrick Tranouez 외

The EXO-POPP project aims to establish a comprehensive database comprising 300,000 marriage records from Paris and its suburbs, spanning the years 1880 to 1940, which are preserved in over 130,000 scans of double pages. …

Handwritten Text Recognition

Neural News Recommendation with Heterogeneous User Behavior

2019-11-01 · IJCNLP 2019 11 · Chuhan Wu, Fangzhao Wu, Mingxiao An, Tao Qi 외

News recommendation is important for online news platforms to help users find interested news and alleviate information overload. Existing news recommendation methods usually rely on the news click history to model user …

MULTI-VIEW LEARNINGNews Recommendation

Extraction of Relevant Images for Boilerplate Removal in Web Browsers

2019-12-17 · Joy Bose

Boilerplate refers to unwanted and repeated parts of a webpage (such as ads or table of contents) that distracts the user from reading the core content of the webpage, such as a news article. Accurate detection and remov…

Processing topical queries on images of historical newspaper pages

2020-02-20 · José E. B. Maia, Gildácio J. de A. Sá

Historical newspapers are a source of research for the human and social sciences. However, these image collections are difficult to read by machine due to the low quality of the print, the lack of standardization of the …

Retrieval