paper-with-me

Papers

WARC-DL: Scalable Web Archive Processing for Deep Learning

2022-09-25 · Niklas Deckers, Martin Potthast

Web archives have grown to petabytes. In addition to providing invaluable background knowledge on many social and cultural developments over the last 30 years, they also provide vast amounts of training data for machine learning. To benefit from recent developments in Deep Learning, the use of web archives requires a scalable solution for their processing that supports inference with and training of neural networks. To date, there is no publicly available library for processing web archives in this way, and some existing applications use workarounds. This paper presents WARC-DL, a deep learning-enabled pipeline for web archive processing that scales to petabytes.

📄 PDF Abstract BibTeX arXiv:2209.12299

Code (1)

chatnoir-eu/chatnoir-warc-dl 공식 구현 tf

Tasks

Deep Learning

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

FastWARC: Optimizing Large-Scale Web Archive Analytics

2021-11-22 · Janek Bevendorff, Martin Potthast, Benno Stein

Web search and other large-scale web data analytics rely on processing archives of web pages stored in a standardized and efficient format. Since its introduction in 2008, the IIPC's Web ARCive (WARC) format has become t…

WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions

2025-10-10 · Sanjari Srivastava, Gang Li, Cheng Chang, Rishu Garg 외 arxiv

Training web agents to navigate complex, real-world websites requires them to master $\textit{subtasks}$ - short-horizon interactions on multiple UI components (e.g., choosing the correct date in a date picker, or scroll…

Reinforcement Learning

Web Archives Metadata Generation with GPT-4o: Challenges and Insights

2024-11-08 · Ashwin Nair, Zhen Rong Goh, Tianrui Liu, Abigail Yongping Huang

Current metadata creation for web archives is time consuming and costly due to reliance on human effort. This paper explores the use of gpt-4o for metadata generation within the Web Archive Singapore, focusing on scalabi…

Prompt Engineering

Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets

2025-11-22 · Gowtham, Sai Rupesh, Sanjay Kumar, Saravanan 외 arxiv

High-quality training data is fundamental to large language model (LLM) performance, yet existing preprocessing pipelines often struggle to effectively remove noise and unstructured content from web-scale corpora. This p…

Scalable Metric Learning via Weighted Approximate Rank Component Analysis

2016-03-01 · Cijo Jose, Francois Fleuret

We are interested in the large-scale learning of Mahalanobis distances, with a particular focus on person re-identification. We propose a metric learning formulation called Weighted Approximate Rank Component Analysis …

Metric LearningPerson Re-IdentificationStochastic Optimization