FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web Documents
Extracting structured data from HTML documents is a long-studied problem with a broad range of applications like augmenting knowledge bases, supporting faceted search, and providing domain-specific experiences for key verticals like shopping and movies. Previous approaches have either required a small number of examples for each target site or relied on carefully handcrafted heuristics built over visual renderings of websites. In this paper, we present a novel two-stage neural approach, named FreeDOM, which overcomes both these limitations. The first stage learns a representation for each DOM node in the page by combining both the text and markup information. The second stage captures longer range distance and semantic relatedness using a relational neural network. By combining these stages, FreeDOM is able to generalize to unseen sites after training on a small number of seed sites from that vertical without requiring expensive hand-crafted features over visual renderings of the page. Through experiments on a public dataset with 8 different verticals, we show that FreeDOM beats the previous state of the art by nearly 3.7 F1 points on average without requiring features over rendered pages or expensive hand-crafted features.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Label-Efficient Self-Training for Attribute Extraction from Semi-Structured Web Documents
Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely…
AttributeAttribute ExtractionKnowledge Base ConstructionSimplified DOM Trees for Transferable Attribute Extraction from the Web
There has been a steady need to precisely extract structured knowledge from the web (i.e. HTML documents). Given a web page, extracting a structured object along with various attributes of interest (e.g. price, publisher…
AttributeAttribute ExtractionFeature EngineeringKnowledge Base ConstructionA Multilingual Information Extraction Pipeline for Investigative Journalism
We introduce an advanced information extraction pipeline to automatically process very large collections of unstructured textual data for the purpose of investigative journalism. The pipeline serves as a new input proces…
Entity Extraction using GANNew/s/leak 2.0 - Multilingual Information Extraction and Visualization for Investigative Journalism
Investigative journalism in recent years is confronted with two major challenges: 1) vast amounts of unstructured data originating from large text collections such as leaks or answers to Freedom of Information requests, …
Efficient ExplorationUnderstanding 6G through Language Models: A Case Study on LLM-aided Structured Entity Extraction in Telecom Domain
Knowledge understanding is a foundational part of envisioned 6G networks to advance network intelligence and AI-native network architectures. In this paradigm, information extraction plays a pivotal role in transforming …
AttributeDecoder