paper-with-me

Papers

BUbiNG: Massive Crawling for the Masses

2016-01-26 · Boldi Paolo, Marino Andrea, Santini Massimo, Vigna Sebastiano

Although web crawlers have been around for twenty years by now, there is virtually no freely available, opensource crawling software that guarantees high throughput, overcomes the limits of single-machine systems and at the same time scales linearly with the amount of resources available. This paper aims at filling this gap, through the description of BUbiNG, our next-generation web crawler built upon the authors' experience with UbiCrawler [Boldi et al. 2004] and on the last ten years of research on the topic. BUbiNG is an opensource Java fully distributed crawler; a single BUbiNG agent, using sizeable hardware, can crawl several thousands pages per second respecting strict politeness constraints, both host- and IP-based. Unlike existing open-source distributed crawlers that rely on batch techniques (like MapReduce), BUbiNG job distribution is based on modern high-speed protocols so to achieve very high throughput.

📄 PDF Abstract BibTeX arXiv:1601.06919

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

287,872 Supermassive Black Holes Masses: Deep Learning Approaching Reverberation Mapping Accuracy

2025-12-04 · Yuhao Lu, HengJian SiTu, Jie Li, Yixuan Li 외 arxiv

We present a population-scale catalogue of 287,872 supermassive black hole masses with high accuracy. Using a deep encoder-decoder network trained on optical spectra with reverberation-mapping (RM) based labels of 849 qu…

Smart Bilingual Focused Crawling of Parallel Documents

2024-05-23 · Cristian García-Romero, Miquel Esplà-Gomis, Felipe Sánchez-Martínez

Crawling parallel texts $\unicode{x2014}$texts that are mutual translations$\unicode{x2014}$ from the Internet is usually done following a brute-force approach: documents are massively downloaded in an unguided process, …

MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages

2022-06-01 · EAMT 2022 6 · Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero 외

We introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual a…

Comparative analysis of various web crawler algorithms

2023-06-21 · Nithin T K, Chandana S, Barani G, Chavva Dharani 외

This presentation focuses on the importance of web crawling and page ranking algorithms in dealing with the massive amount of data present on the World Wide Web. As the web continues to grow exponentially, efficient sear…

Information RetrievalNavigateRetrieval

esCorpius: A Massive Spanish Crawling Corpus

2022-06-30 · Asier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé, David Griol 외

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack o…

Language Modelling