paper-with-me

홈 › Papers

Evaluating the Usefulness of Sentiment Information for Focused Crawlers

2013-09-27 · Tianjun Fu, Ahmed Abbasi, Daniel Zeng, Hsinchun Chen

Despite the prevalence of sentiment-related content on the Web, there has been limited work on focused crawlers capable of effectively collecting such content. In this study, we evaluated the efficacy of using sentiment-related information for enhanced focused crawling of opinion-rich web content regarding a particular topic. We also assessed the impact of using sentiment-labeled web graphs to further improve collection accuracy. Experimental results on a large test bed encompassing over half a million web pages revealed that focused crawlers utilizing sentiment information as well as sentiment-labeled web graphs are capable of gathering more holistic collections of opinion-related content regarding a particular topic. The results have important implications for business and marketing intelligence gathering efforts in the Web 2.0 era.

📄 PDF Abstract BibTeX arXiv:1309.7270

Code (0)

등록된 구현이 없습니다.

Tasks

Marketing

Similar Papers 제목 키워드 기반

The iCrawl Wizard -- Supporting Interactive Focused Crawl Specification

2016-12-19 · Gossen Gerhard, Demidova Elena, Risse Thomas

Collections of Web documents about specific topics are needed for many areas of current research. Focused crawling enables the creation of such collections on demand. Current focused crawlers require the user to manually…

Comparing the Quality of Focused Crawlers and of the Translation Resources Obtained from them

2014-05-01 · LREC 2014 5 · Bruno Laranjeira, Viviane Moreira, Aline Villavicencio, Carlos Ramisch 외

Comparable corpora have been used as an alternative for parallel corpora as resources for computational tasks that involve domain-specific natural language processing. One way to gather documents related to a specific to…

Machine TranslationTranslation

PriPA: A Tool for Privacy-Preserving Analytics of Linguistic Data

2022-06-01 · LEGAL (LREC) 2022 6 · Jeremie Clos, Emma McClaughlin, Pepita Barnard, Elena Nichele 외

The days of large amorphous corpora collected with armies of Web crawlers and stored indefinitely are, or should be, coming to an end. There is a wealth of hidden linguistic information that is increasingly difficult to …

Privacy Preserving

iCrawl: Improving the Freshness of Web Collections by Integrating Social Web and Focused Web Crawling

2016-12-19 · Gossen Gerhard, Demidova Elena, Risse Thomas

Researchers in the Digital Humanities and journalists need to monitor, collect and analyze fresh online content regarding current events such as the Ebola outbreak or the Ukraine crisis on demand. However, existing focus…

ANTUSD: A Large Chinese Sentiment Dictionary

2016-05-01 · LREC 2016 5 · Shih-Ming Wang, Lun-Wei Ku

This paper introduces the augmented NTU sentiment dictionary, abbreviated as ANTUSD, which is constructed by collecting sentiment stats of words in several sentiment annotation work. A total of 26,021 words were collecte…

General Classification