paper-with-me

홈 › Papers

Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method

2026-06-06 · Kyoungmin Kim, Martin Catheland, Anastasia Ailamaki arxiv

Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of LLM-based data processing. Calling the LLM on every document (the oracle) is prohibitive, so cascades pair the oracle with a fast proxy. As deployed today, they leave four limitations on the table. (1) Each cascade family - model-free clustering, prebuilt small-LLM proxies, online-trained proxies - commits to a single representation and pipeline, and wins on only a narrow query regime. (2) The strongest online proxy invests in a custom training scheme on a bi-encoder over dense embeddings, missing the token-level evidence richer predicates require. (3) The proxy is trained against binary yes/no labels, wasting the LLM's per-document confidence at the boundary documents it most needs to learn. (4) Existing calibrations add a uniform safety margin, conflating genuine proxy uncertainty with small-sample noise and inflating cascade cost. We address these by (1) composing families adaptively - model-free clustering first, online proxy only when needed, with oracle calls shared across phases; (2) replacing the cosine bi-encoder with a hybrid of off-the-shelf token-aware models; (3) training the proxy with the oracle's per-document confidence as a soft label; and (4) a calibration that adds the safety margin only where the labeled sample is sparse. We are also the first to use the oracle's per-document confidence for three purposes: a query-level difficulty compass, a lower bound on the minimum oracle calls any proxy-based cascade can make, and the proxy's soft training label. At a 90% accuracy target on three 10K-document corpora, our methods are 1.6-2.0x faster than the best prior method per corpus and meet the target on 95% of queries; the BER-derived lower bound indicates a further ~4-20x of headroom for future work.

📄 PDF Abstract BibTeX arXiv:2606.08090

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fast Semantic Image Segmentation with High Order Context and Guided Filtering

2016-05-13 · Falong Shen, Gang Zeng

This paper describes a fast and accurate semantic image segmentation approach that encodes not only the discriminative features from deep neural networks, but also the high-order context compatibility among adjacent obje…

Image SegmentationSemantic SegmentationVocal Bursts Intensity Prediction

Sparsity-Aware Robust Normalized Subband Adaptive Filtering algorithms based on Alternating Optimization

2022-05-15 · Yi Yu, Zongxin Huang, Hongsen He, Yuriy Zakharov 외

This paper proposes a unified sparsity-aware robust normalized subband adaptive filtering (SA-RNSAF) algorithm for identification of sparse systems under impulsive noise. The proposed SA-RNSAF algorithm generalizes diffe…

Faster-TAD: Towards Temporal Action Detection with Proposal Generation and Classification in a Unified Network

2022-04-06 · Shimin Chen, Chen Chen, Wei Li, Xunqiang Tao 외

Temporal action detection (TAD) aims to detect the semantic labels and boundaries of action instances in untrimmed videos. Current mainstream approaches are multi-step solutions, which fall short in efficiency and flexib…

Action DetectionAction Spotting

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

2025-03-12 · Xiaowei Bi, Zheyuan Xu

Long Video Question Answering (LVQA) is challenging due to the need for temporal reasoning and large-scale multimodal data processing. Existing methods struggle with retrieving cross-modal information from long videos, e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Information RetrievalQuestion Answering+8

Efficient Model-Based Collaborative Filtering with Fast Adaptive PCA

2020-09-04 · Xiangyun Ding, Wenjian Yu, Yuyang Xie, Shenghua Liu

A model-based collaborative filtering (CF) approach utilizing fast adaptive randomized singular value decomposition (SVD) is proposed for the matrix completion problem in recommender system. Firstly, a fast adaptive PCA …

Collaborative FilteringMatrix CompletionRecommendation Systems