paper-with-me

Papers

Can We Find Documents in Web Archives without Knowing their Contents?

2017-01-14 · Vo Khoi Duy, Tran Tuan, Nguyen Tu Ngoc, Zhu Xiaofei, Nejdl Wolfgang

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access and exploration mechanisms. Ideal search and rank- ing methods must be robust to the high redundancy and the temporal noise of contents, as well as scalable to the huge amount of data archived. Despite several attempts in Web archive search, facilitating access to Web archive still remains a challenging problem. In this work, we conduct a first analysis on different ranking strategies that exploit evidences from metadata instead of the full content of documents. We perform a first study to compare the usefulness of non-content evidences to Web archive search, where the evidences are mined from the metadata of file headers, links and URL strings only. Based on these findings, we propose a simple yet surprisingly effective learning model that combines multiple evidences to distinguish "good" from "bad" search results. We conduct empirical experiments quantitatively as well as qualitatively to confirm the validity of our proposed method, as a first step towards better ranking in Web archives taking meta- data into account.

📄 PDF Abstract BibTeX arXiv:1701.03942

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Archives: Towards New Research Interfaces Relying on the Semantic Annotation of Documents

2024-03-28 · Nicolas Gutehrlé, Iana Atanassova

The digitisation campaigns carried out by libraries and archives in recent years have facilitated access to documents in their collections. However, exploring and exploiting these documents remain difficult tasks due to …

Handwriting Classification for the Analysis of Art-Historical Documents

2020-11-04 · Christian Bartz, Hendrik Rätz, Christoph Meinel

Digitized archives contain and preserve the knowledge of generations of scholars in millions of documents. The size of these archives calls for automatic analysis since a manual analysis by specialists is often too expen…

ClassificationGeneral ClassificationOptical Character Recognition (OCR)text-classification+1

SIMARA: a database for key-value information extraction from full pages

2023-04-26 · Solène Tarride, Mélodie Boillet, Jean-François Moufflet, Christopher Kermorvant

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwrit…

Handwriting RecognitionHandwritten Text RecognitionKey Information ExtractionNamed Entity Recognition (NER)

Towards a Ranking Model for Semantic Layers over Digital Archives

2018-10-23 · Fafalios Pavlos, Kasturia Vaibhav, Nejdl Wolfgang

Archived collections of documents (like newspaper archives) serve as important information sources for historians, journalists, sociologists and other interested parties. Semantic Layers over such digital archives allow …

Archaeology in the Digital Age: From Paper to Databases

2015-07-08 · Frédérique Mélanie-Becquet, Johan Ferguth, Katherine Gruel, Thierry Poibeau

Research units in archaeology often manage large and precious archives containing various documents, including reports on fieldwork, scholarly studies and reference books. These archives are of course invaluable, recordi…