paper-with-me

홈 › Papers

Extracting News Web Page Creation Time with DCTFinder

2014-05-01 · LREC 2014 5 · Xavier Tannier

Web pages do not offer reliable metadata concerning their creation date and time. However, getting the document creation time is a necessary step for allowing to apply temporal normalization systems to web pages. In this paper, we present DCTFinder, a system that parses a web page and extracts from its content the title and the creation date of this web page. DCTFinder combines heuristic title detection, supervised learning with Conditional Random Fields (CRFs) for document date extraction, and rule-based creation time recognition. Using such a system allows further deep and efficient temporal analysis of web pages. Evaluation on three corpora of English and French web pages indicates that the tool can extract document creation times with reasonably high accuracy (between 87 and 92{\textbackslash}{\%}). DCTFinder is made freely available on http://sourceforge.net/projects/dctfinder/, as well as all resources (vocabulary and annotated documents) built for training and evaluating the system in English and French, and the English trained model itself.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Document SummarizationInformation RetrievalMulti-Document SummarizationQuestion Answering

Similar Papers 제목 키워드 기반

How much is Wikipedia Lagging Behind News?

2017-03-30 · Fetahu Besnik, Anand Abhijit, Anand Avishek

Wikipedia, rich in entities and events, is an invaluable resource for various knowledge harvesting, extraction and mining tasks. Numerous resources like DBpedia, YAGO and other knowledge bases are based on extracting ent…

Articles

Multilingual Attribute Extraction from News Web Pages

2025-02-04 · Pavel Bedrin, Maksim Varlamov, Alexander Yatskov

This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high efficacy in extracting information from semi-s…

AttributeAttribute Extraction

Data Set for Stance and Sentiment Analysis from User Comments on Croatian News

2019-08-01 · WS 2019 8 · Mihaela Bo{\v{s}}njak, Mladen Karan

Nowadays it is becoming more important than ever to find new ways of extracting useful information from the evergrowing amount of user-generated data available online. In this paper, we describe the creation of a data se…

ArticlesBIG-bench Machine LearningSentiment Analysis

Experimenting AI Technologies for Disinformation Combat: the IDMO Project

2023-10-17 · Lorenzo Canale, Alberto Messina

The Italian Digital Media Observatory (IDMO) project, part of a European initiative, focuses on countering disinformation and fake news. This report outlines contributions from Rai-CRITS to the project, including: (i) th…

Natural Language Inference

Analysis of User Dwell Time by Category in News Application

2019-08-23 · Yoshifumi Seki, Mitsuo Yoshida

Dwell time indicates how long a user looked at a page, and this is used especially in fields where ratings from users such as search engines, recommender systems, and advertisements are important. Despite the importance …

Recommendation Systems