paper-with-me

홈 › Papers

OSCAR: Online Soft Compression And Reranking

2025-03-17 · Maxime Louis, Thibault Formal, Hervé Dejean, Stéphane Clinchant

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge, leading to improved accuracy and relevance. However, scaling RAG pipelines remains computationally expensive as retrieval sizes grow. To address this, we introduce OSCAR, a novel query-dependent online soft compression method that reduces computational overhead while preserving performance. Unlike traditional hard compression methods, which shorten retrieved texts, or soft compression approaches, which map documents to continuous embeddings offline, OSCAR dynamically compresses retrieved information at inference time, eliminating storage overhead and enabling higher compression rates. Additionally, we extend OSCAR to simultaneously perform reranking, further optimizing the efficiency of the RAG pipeline. Our experiments demonstrate state-of-the-art performance with a 2-5x speed-up in inference and minimal to no loss in accuracy for LLMs ranging from 1B to 24B parameters. The models are available at: https://huggingface.co/collections/naver/oscar-67d446a8e3a2551f57464295.

📄 PDF Abstract BibTeX arXiv:2504.07109

Code (0)

등록된 구현이 없습니다.

Tasks

RAGRerankingRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Adam 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

OSCAR: One-Step Diffusion Codec for Image Compression Across Multiple Bit-rates

2025-05-22 · Jinpei Guo, Yifei Ji, Zheng Chen, Kai Liu 외

Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising fro…

DenoisingImage Compression

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

2026-05-19 · Zunhai Su, Rui Yang, Chao Zhang, Yaxiu Liu 외 arxiv

The rapid advancement toward long-context reasoning and multi-modal intelligence has made the memory footprint of the Key-Value (KV) cache a dominant memory bottleneck for efficient deployment. While the established per-…

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

2026-08-08 · Gunjan Paul, Senthil Palanisamy, Satpal Singh Rathore, Pratyush Kumar Patnaik 외 hf

We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis…

OSCAR: Optimization-Steered Agentic Planning for Composed Image Retrieval

2026-02-09 · Teng Wang, Rong Shan, Jianghao Lin, Junjie Wu 외 arxiv

Composed image retrieval (CIR) requires complex reasoning over heterogeneous visual and textual constraints. Existing approaches largely fall into two paradigms: unified embedding retrieval, which suffers from single-mod…

Image Retrieval

OSCAR-Net: Object-centric Scene Graph Attention for Image Attribution

2021-08-07 · ICCV 2021 10 · Eric Nguyen, Tu Bui, Vishy Swaminathan, John Collomosse

Images tell powerful stories but cannot always be trusted. Matching images back to trusted sources (attribution) enables users to make a more informed judgment of the images they encounter online. We propose a robust ima…

Contrastive LearningGraph AttentionImage Attribution