paper-with-me

홈 › Papers

The iCrawl Wizard -- Supporting Interactive Focused Crawl Specification

2016-12-19 · Gossen Gerhard, Demidova Elena, Risse Thomas

Collections of Web documents about specific topics are needed for many areas of current research. Focused crawling enables the creation of such collections on demand. Current focused crawlers require the user to manually specify starting points for the crawl (seed URLs). These are also used to describe the expected topic of the collection. The choice of seed URLs influences the quality of the resulting collection and requires a lot of expertise. In this demonstration we present the iCrawl Wizard, a tool that assists users in defining focused crawls efficiently and semi-automatically. Our tool uses major search engines and Social Media APIs as well as information extraction techniques to find seed URLs and a semantic description of the crawl intent. Using the iCrawl Wizard even non-expert users can create semantic specifications for focused crawlers interactively and efficiently.

📄 PDF Abstract BibTeX arXiv:1612.06162

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Wizard Computer vision is an interesting tool for animal behavior monitoring, mainly because it limits animal handling and it can be used to record various traits using only one sensor.…

Similar Papers 제목 키워드 기반

iCrawl: Improving the Freshness of Web Collections by Integrating Social Web and Focused Web Crawling

2016-12-19 · Gossen Gerhard, Demidova Elena, Risse Thomas

Researchers in the Digital Humanities and journalists need to monitor, collect and analyze fresh online content regarding current events such as the Ebola outbreak or the Ukraine crisis on demand. However, existing focus…

BUbiNG: Massive Crawling for the Masses

2016-01-26 · Boldi Paolo, Marino Andrea, Santini Massimo, Vigna Sebastiano

Although web crawlers have been around for twenty years by now, there is virtually no freely available, opensource crawling software that guarantees high throughput, overcomes the limits of single-machine systems and at …

RecWizard: A Toolkit for Conversational Recommendation with Modular, Portable Models and Interactive User Interface

2024-02-23 · Zeyuan Zhang, Tanmay Laud, Zihang He, Xiaojie Chen 외

We present a new Python toolkit called RecWizard for Conversational Recommender Systems (CRS). RecWizard offers support for development of models and interactive user interface, drawing from the best practices of the Hug…

Conversational RecommendationRecommendation Systems

Domain-Specific Corpus Expansion with Focused Webcrawling

2016-05-01 · LREC 2016 5 · Steffen Remus, Chris Biemann

This work presents a straightforward method for extending or creating in-domain web corpora by focused webcrawling. The focused webcrawler uses statistical N-gram language models to estimate the relatedness of documents …

EasySpider: A No-Code Visual System for Crawling the Web

2023-04-30 · ACM The Web Conference 2023 4 · Naibo Wang, Wenjie Feng, Jianwei Yin, See-Kiong Ng

The web is a treasure trove for data that is increasingly used by computer scientists for building large machine learning models as well as non-computer scientists for social studies or marketing analyses. As such, web-c…

Data IntegrationMarketing