paper-with-me

홈 › Papers

Boilerplate Detection via Semantic Classification of TextBlocks

2022-03-09 · Hao Zhang, Jie Wang

We present a hierarchical neural network model called SemText to detect HTML boilerplate based on a novel semantic representation of HTML tags, class names, and text blocks. We train SemText on three published datasets of news webpages and fine-tune it using a small number of development data in CleanEval and GoogleTrends-2017. We show that SemText achieves the state-of-the-art accuracy on these datasets. We then demonstrate the robustness of SemText by showing that it also detects boilerplate effectively on out-of-domain community-based question-answer webpages.

📄 PDF Abstract BibTeX arXiv:2203.04467

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Similar Papers 제목 키워드 기반

Extraction of Relevant Images for Boilerplate Removal in Web Browsers

2019-12-17 · Joy Bose

Boilerplate refers to unwanted and repeated parts of a webpage (such as ads or table of contents) that distracts the user from reading the core content of the webpage, such as a news article. Accurate detection and remov…

M-PACT: An Open Source Platform for Repeatable Activity Classification Research

2018-04-16 · Eric Hofesmann, Madan Ravi Ganesh, Jason J. Corso

There are many hurdles that prevent the replication of existing work which hinders the development of new activity classification models. These hurdles include switching between multiple deep learning libraries and the d…

Action ClassificationActivity RecognitionClassificationGeneral Classification

Semi-Supervised Method using Gaussian Random Fields for Boilerplate Removal in Web Browsers

2019-11-08 · Joy Bose, Sumanta Mukherjee

Boilerplate removal refers to the problem of removing noisy content from a webpage such as ads and extracting relevant content that can be used by various services. This can be useful in several features in web browsers …

BlockingTranslation

On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective

2026-06-01 · Yikai Guo, Bin Wang, Xilai Fan, Wenjun Ke 외 arxiv

AI-generated text increasingly blends with human writing, raising practical risks such as misinformation, academic misuse, and corpora contamination. While statistical detectors are appealing for efficiency and generaliz…

Text Detection

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

2026-07-02 · Xinyi Fang, Kejian Tong, Jiabei Liu, Tao Ning 외 arxiv

Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attent…