paper-with-me

홈 › Papers

WebAnnotator, an Annotation Tool for Web Pages

2012-05-01 · LREC 2012 5 · Xavier Tannier

This article presents WebAnnotator, a new tool for annotating Web pages. WebAnnotator is implemented as a Firefox extension, allowing annotation of both offline and inline pages. The HTML rendering fully preserved and all annotations consist in new HTML spans with specific styles. WebAnnotator provides an easy and general-purpose framework and is made available under CeCILL free license (close to GNU GPL), so that use and further contributions are made simple. All parts of an HTML document can be annotated: text, images, videos, tables, menus, etc. The annotations are created by simply selecting a part of the document and clicking on the relevant type and subtypes. The annotated elements are then highlighted in a specific color. Annotation schemas can be defined by the user by creating a simple DTD representing the types and subtypes that must be highlighted. Finally, annotations can be saved (HTML with highlighted parts of documents) or exported (in a machine-readable format).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalInformation Retrieval

Similar Papers 제목 키워드 기반

Tag-Pag: A Dedicated Tool for Systematic Web Page Annotations

2025-02-22 · Anton Pogrebnjak, Julian Schelb, Andreas Spitz, Celina Kacperski 외

Tag-Pag is an application designed to simplify the categorization of web pages, a task increasingly common for researchers who scrape web pages to analyze individuals' browsing patterns or train machine learning classifi…

TAG

A Co-guided Neural Network for Person Name Recognition in Academic Homepages

2019-05-14 · Anonymous

Academic homepages are important channels for learning researchers' profiles. Knowing the person names in academic homepages is essential to the extraction of other entities such as contacts, publications, and biography.…

NER

Visual-Aware Representation of Web Pages for Machine Learning Applications

2026-08-19 · Radek Burget, Radek Hranický arxiv

Applying machine learning to web pages is challenging due to the need to interpret HTML together with associated resources and perform rendering to obtain a meaningful visual and layout-aware representation. As a result,…

Efficacy of AI RAG Tools for Complex Information Extraction and Data Annotation Tasks: A Case Study Using Banks Public Disclosures

2025-07-28 · Nicholas Botti, Flora Haberkorn, Charlotte Hoopes, Shaun Khan arxiv

We utilize a within-subjects design with randomized task assignments to understand the effectiveness of using an AI retrieval augmented generation (RAG) tool to assist analysts with an information extraction and data ann…

Information Extraction

Inforex -- a web-based tool for text corpus management and semantic annotation

2012-05-01 · LREC 2012 5 · Micha{\l} Marci{\'n}czuk, Jan Koco{\'n}, Bartosz Broda

The aim of this paper is to present a system for semantic text annotation called Inforex. Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Name…

ManagementNamed Entity Recognition (NER)SentenceSentence segmentation+3