paper-with-me

홈 › Papers

"Piaf" vs "Adele": classifying encyclopedic queries using automatically labeled training data

2015-11-30 · Saleiro Pedro, Sarmento Luís

Encyclopedic queries express the intent of obtaining information typically available in encyclopedias, such as biographical, geographical or historical facts. In this paper, we train a classifier for detecting the encyclopedic intent of web queries. For training such a classifier, we automatically label training data from raw query logs. We use click-through data to select positive examples of encyclopedic queries as those queries that mostly lead to Wikipedia articles. We investigated a large set of features that can be generated to describe the input query. These features include both term-specific patterns as well as query projections on knowledge bases items (e.g. Freebase). Results show that using these feature sets it is possible to achieve an F1 score above 87%, competing with a Google-based baseline, which uses a much wider set of signals to boost the ranking of Wikipedia for potential encyclopedic queries. The results also show that both query projections on Wikipedia article titles and Freebase entity match represent the most relevant groups of features. When the training set contains frequent positive examples (i.e rare queries are excluded) results tend to improve.

📄 PDF Abstract BibTeX arXiv:1511.09290

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Enriching Ontologies with Encyclopedic Background Knowledge for Document Indexing

2016-03-21 · Posch Lisa

The rapidly increasing number of scientific documents available publicly on the Internet creates the challenge of efficiently organizing and indexing these documents. Due to the time consuming and tedious nature of manua…

BIG-bench Machine Learning

PIAFusion: A progressive infrared and visible image fusion network based on illumination aware

2022-07-01 · Information Fusion 2022 7 · LinfengTang, JitengYuan, HaoZhang, XingyuJiang 외

Infrared and visible image fusion aims to synthesize a single fused image containing salient targets and abundant texture details even under extreme illumination conditions. However, existing image fusion algorithms fail…

Infrared And Visible Image FusionSemantic Segmentation

Sustainable Modular Debiasing of Language Models

2021-09-08 · Findings (EMNLP) 2021 11 · Anne Lauscher, Tobias Lüken, Goran Glavaš

Unfair stereotypical biases (e.g., gender, racial, or religious biases) encoded in modern pretrained language models (PLMs) have negative ethical implications for widespread adoption of state-of-the-art language technolo…

FairnessLanguage ModelingLanguage Modelling

Multistain Pretraining for Slide Representation Learning in Pathology

2024-08-05 · Guillaume Jaume, Anurag Vaidya, Andrew Zhang, Andrew H. Song 외

Developing self-supervised learning (SSL) models that can learn universal and transferable representations of H&E gigapixel whole-slide images (WSIs) is becoming increasingly valuable in computational pathology. These mo…

Representation LearningSelf-Supervised Learningwhole slide images

The ADELE Corpus of Dyadic Social Text Conversations:Dialog Act Annotation with ISO 24617-2

2018-05-01 · LREC 2018 5 · Emer Gilmartin, Christian Saam, Brendan Spillane, Maria O{'}Reilly 외
Chatbot