paper-with-me

홈 › Papers

Automatic Query Optimization for Retrieving Traffic Tweets

2020-06-21 · Emory Hufbauer, Hana Khamfroush

Twitter, like many social media and data brokering companies, makes their data available through a search API (application programming interface). In addition to filtering results by date and location, researchers can search for tweets with specific content with a boolean text query, using {\it AND}, {\it OR}, and {\it NOT} operators to select the combinations of phrases which must, or must not, appear in matching tweets. This boolean text search system is not at all unique to Twitter and is found in many different contexts, including academic, legal, and medical databases, however it is stretched to its limits in Twitter's use case because of the relative volume and brevity of tweets. In addition, the semi-automated use of such systems was well studied under the topic of Information Retrieval during the 1980s and 1990s, however the study of such systems has greatly declined since that time. As such, we propose updated methods for automatically selecting and refining complex boolean search queries that can isolate relevant results with greater specificity and completeness. Furthermore, we present preliminary results of using an optimized query to collect a sample of traffic-incident-related tweets, along with the results of manually classifying and analyzing them.

📄 PDF Abstract BibTeX arXiv:2006.11887

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrievalSpecificity

Similar Papers 제목 키워드 기반

Fake News Data Collection and Classification: Iterative Query Selection for Opaque Search Engines with Pseudo Relevance Feedback

2020-12-23 · Aviad Elyashar, Maor Reuben, Rami Puzis

Retrieving information from an online search engine, is the first and most important step in many data mining tasks. Most of the search engines currently available on the web, including all social media platforms, are bl…

Fake News Detection

HITSZ-ICRC: A Report for SMM4H Shared Task 2020-Automatic Classification of Medications and Adverse Effect in Tweets

2020-12-01 · SMM4H (COLING) 2020 12 · Xiaoyu Zhao, Ying Xiong, Buzhou Tang

This is the system description of the Harbin Institute of Technology Shenzhen (HITSZ) team for the first and second subtasks of the fifth Social Media Mining for Health Applications (SMM4H) shared task in 2020. The first…

ClassificationTask 2

A Web Scraping Methodology for Bypassing Twitter API Restrictions

2018-03-27 · Hernandez-Suarez A., Sanchez-Perez G., Toscano-Medina K., Martinez-Hernandez V. 외

Retrieving information from social networks is the first and primordial step many data analysis fields such as Natural Language Processing, Sentiment Analysis and Machine Learning. Important data science tasks relay on h…

BIG-bench Machine LearningSentiment Analysis

Twitter-based traffic information system based on vector representations for words

2018-12-04 · Sina Dabiri, Kevin Heaslip

Recently, researchers have shown an increased interest in harnessing Twitter data for dynamic monitoring of traffic conditions. Bag-of-words representation is a common method in literature for tweet modeling and retrievi…

Semantic SimilaritySemantic Textual Similarity

Smart Crawling: A New Approach toward Focus Crawling from Twitter

2021-10-08 · Ahmad Khazaie, Nacéra Bennacer Seghouani, Francesca Bugiotti

Twitter is a social network that offers a rich and interesting source of information challenging to retrieve and analyze. Twitter data can be accessed using a REST API. The available operations allow retrieving tweets on…