paper-with-me

홈 › Papers

FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web Documents

2020-10-21 · Bill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep Tata

Extracting structured data from HTML documents is a long-studied problem with a broad range of applications like augmenting knowledge bases, supporting faceted search, and providing domain-specific experiences for key verticals like shopping and movies. Previous approaches have either required a small number of examples for each target site or relied on carefully handcrafted heuristics built over visual renderings of websites. In this paper, we present a novel two-stage neural approach, named FreeDOM, which overcomes both these limitations. The first stage learns a representation for each DOM node in the page by combining both the text and markup information. The second stage captures longer range distance and semantic relatedness using a relational neural network. By combining these stages, FreeDOM is able to generalize to unseen sites after training on a small number of seed sites from that vertical without requiring expensive hand-crafted features over visual renderings of the page. Through experiments on a public dataset with 8 different verticals, we show that FreeDOM beats the previous state of the art by nearly 3.7 F1 points on average without requiring features over rendered pages or expensive hand-crafted features.

📄 PDF Abstract BibTeX arXiv:2010.10755

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Label-Efficient Self-Training for Attribute Extraction from Semi-Structured Web Documents

2022-08-27 · Ritesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant Shiralkar

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely…

AttributeAttribute ExtractionKnowledge Base Construction

Simplified DOM Trees for Transferable Attribute Extraction from the Web

2021-01-07 · Yichao Zhou, Ying Sheng, Nguyen Vo, Nick Edmonds 외

There has been a steady need to precisely extract structured knowledge from the web (i.e. HTML documents). Given a web page, extracting a structured object along with various attributes of interest (e.g. price, publisher…

AttributeAttribute ExtractionFeature EngineeringKnowledge Base Construction

A Multilingual Information Extraction Pipeline for Investigative Journalism

2018-09-01 · EMNLP 2018 11 · Gregor Wiedemann, Seid Muhie Yimam, Chris Biemann

We introduce an advanced information extraction pipeline to automatically process very large collections of unstructured textual data for the purpose of investigative journalism. The pipeline serves as a new input proces…

Entity Extraction using GAN

New/s/leak 2.0 - Multilingual Information Extraction and Visualization for Investigative Journalism

2018-07-13 · Gregor Wiedemann, Seid Muhie Yimam, Chris Biemann

Investigative journalism in recent years is confronted with two major challenges: 1) vast amounts of unstructured data originating from large text collections such as leaks or answers to Freedom of Information requests, …

Efficient Exploration

Understanding 6G through Language Models: A Case Study on LLM-aided Structured Entity Extraction in Telecom Domain

2025-05-20 · Ye Yuan, Haolun Wu, Hao Zhou, Xue Liu 외

Knowledge understanding is a foundational part of envisioned 6G networks to advance network intelligence and AI-native network architectures. In this paradigm, information extraction plays a pivotal role in transforming …

AttributeDecoder