paper-with-me

Papers

Pairwise Judgment Formulation for Semantic Embedding Model in Web Search

2024-08-08 · Mengze Hong, Di Jiang, Wailing Ng, Zichang Guo, Chen Jason Zhang

Semantic Embedding Model (SEM), a neural network-based Siamese architecture, is gaining momentum in information retrieval and natural language processing. In order to train SEM in a supervised fashion for Web search, the search engine query log is typically utilized to automatically formulate pairwise judgments as training data. Despite the growing application of semantic embeddings in the search engine industry, little work has been done on formulating effective pairwise judgments for training SEM. In this paper, we make the first in-depth investigation of a wide range of strategies for generating pairwise judgments for SEM. An interesting (perhaps surprising) discovery reveals that the conventional pairwise judgment formulation strategy wildly used in the field of pairwise Learning-to-Rank (LTR) is not necessarily effective for training SEM. Through a large-scale empirical study based on query logs and click-through activities from a major commercial search engine, we demonstrate the effective strategies for SEM and highlight the advantages of a hybrid heuristic (i.e., Clicked > Non-Clicked) in comparison to the atomic heuristics (e.g., Clicked > Skipped) in LTR. We conclude with best practices for training SEM and offer promising insights for future research.

📄 PDF Abstract BibTeX arXiv:2408.04197

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalLearning-To-Rank

Similar Papers 제목 키워드 기반

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

2026-08-19 · Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang 외 arxiv

Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically …

SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding

2025-03-09 · Juhyeon Park, Peter Yongho Kim, Jiook Cha, Shinjae Yoo 외

We present SEED (\textbf{Se}mantic \textbf{E}valuation for Visual Brain \textbf{D}ecoding), a novel metric for evaluating the semantic decoding performance of visual brain decoding models. It integrates three complementa…

Brain DecodingSemantic SimilaritySemantic Textual Similarity

Semantic Data Set Construction from Human Clustering and Spatial Arrangement

2021-03-01 · CL (ACL) 2021 3 · Olga Majewska, Diana McCarthy, Jasper J. F. van den Bosch, Nikolaus Kriegeskorte 외

Abstract Research into representation learning models of lexical semantics usually utilizes some form of intrinsic evaluation to ensure that the learned representations reflect human semantic judgments. Lexical semantic …

ClusteringRepresentation LearningSemantic SimilaritySemantic Textual Similarity+1

Table Integration in Data Lakes Unleashed: Pairwise Integrability Judgment, Integrable Set Discovery, and Multi-Tuple Conflict Resolution

2024-11-30 · Daomin Ji, Hui Luo, Zhifeng Bao, Shane Culpepper

Table integration aims to create a comprehensive table by consolidating tuples containing relevant information. In this work, we investigate the challenge of integrating multiple tables from a data lake, focusing on thre…

Community DetectionContrastive LearningData AugmentationIn-Context Learning

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

2026-08-26 · Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso 외 arxiv

Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles …

Semantic SimilaritySpeech Recognition