paper-with-me

홈 › Papers

EveTAR: Building a Large-Scale Multi-Task Test Collection over Arabic Tweets

2017-08-21 · Hasanain Maram, Suwaileh Reem, Elsayed Tamer, Kutlu Mucahid, Almerekhi Hind

This article introduces a new language-independent approach for creating a large-scale high-quality test collection of tweets that supports multiple information retrieval (IR) tasks without running a shared-task campaign. The adopted approach (demonstrated over Arabic tweets) designs the collection around significant (i.e., popular) events, which enables the development of topics that represent frequent information needs of Twitter users for which rich content exists. That inherently facilitates the support of multiple tasks that generally revolve around events, namely event detection, ad-hoc search, timeline generation, and real-time summarization. The key highlights of the approach include diversifying the judgment pool via interactive search and multiple manually-crafted queries per topic, collecting high-quality annotations via crowd-workers for relevancy and in-house annotators for novelty, filtering out low-agreement topics and inaccessible tweets, and providing multiple subsets of the collection for better availability. Applying our methodology on Arabic tweets resulted in EveTAR , the first freely-available tweet test collection for multiple IR tasks. EveTAR includes a crawl of 355M Arabic tweets and covers 50 significant events for which about 62K tweets were judged with substantial average inter-annotator agreement (Kappa value of 0.71). We demonstrate the usability of EveTAR by evaluating existing algorithms in the respective tasks. Results indicate that the new collection can support reliable ranking of IR systems that is comparable to similar TREC collections, while providing strong baseline results for future studies over Arabic tweets.

📄 PDF Abstract BibTeX arXiv:1708.05517

Code (0)

등록된 구현이 없습니다.

Tasks

Event DetectionInformation RetrievalRetrieval

Similar Papers 제목 키워드 기반

A Web Scale Entity Extraction System

2021-08-27 · Findings (EMNLP) 2021 11 · Xuanting Cai, Quanbin Ma, Pan Li, Jianyu Liu 외

Understanding the semantic meaning of content on the web through the lens of entities and concepts has many practical advantages. However, when building large-scale entity extraction systems, practitioners are facing uni…

Predicting building types and functions at transnational scale

2024-09-15 · Jonas Fill, Michael Eichelbeck, Michael Ebner

Building-specific knowledge such as building type and function information is important for numerous energy applications. However, comprehensive datasets containing this information for individual households are missing …

Graph Neural Network

Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries

2015-07-20 · Yuke Zhu, Ce Zhang, Christopher Ré, Li Fei-Fei

The complexity of the visual world creates significant challenges for comprehensive visual understanding. In spite of recent successes in visual recognition, today's vision systems would still struggle to deal with visua…

Knowledge Base ConstructionRetrieval

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity

2025-09-11 · Jackson Eshbaugh, Chetan Tiwari, Jorge Silveyra arxiv

Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems. However, these model…

Building Extraction at Scale using Convolutional Neural Network: Mapping of the United States

2018-05-23 · Hsiuhan Lexie Yang, Jiangye Yuan, Dalton Lunga, Melanie Laverdiere 외

Establishing up-to-date large scale building maps is essential to understand urban dynamics, such as estimating population, urban planning and many other applications. Although many computer vision tasks has been success…