paper-with-me

Papers

Context Steering: A New Paradigm for Compression-based Embeddings by Synthesizing Relevant Information Features

2025-08-20 · Guillermo Sarasa, Ana Granados, Francisco de Borja Rodríguez arxiv

Compression-based dissimilarities (CD) offer a flexible and domain-agnostic means of measuring similarity by identifying implicit information through redundancies between data objects. However, as similarity features are derived from the data, rather than defined as an input, it often proves difficult to align with the task at hand, particularly in complex clustering or classification settings. To address this issue, we introduce "context steering", a novel methodology that actively guides the feature-shaping process. Instead of passively accepting the emergent data structure (typically a hierarchy derived from clustering CDs), our approach "steers" the process by systematically analyzing how each object influences the relational context within a clustering framework. This process generates a custom-tailored embedding that isolates and amplifies class-distinctive information. We validate this supervised context-steering strategy using Normalized Compression Distance (NCD) and Relative Compression Distance (NRC) combined with hierarchical clustering, and evaluate the learned embeddings through both classification performance and cluster-quality metrics. Experiments on heterogeneous datasets-from text to real-world audio-show that the proposed approach yields robust task-oriented embeddings from compression dissimilarities, moving from traditional transductive uses of distance matrices to an inductive representation that can be applied to unseen data.

📄 PDF Abstract BibTeX arXiv:2508.14780

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction

2026-03-20 · Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu 외 arxiv

While context compression can mitigate the growing inference costs of Large Language Models (LLMs) by shortening contexts, existing methods that specify a target compression ratio or length suffer from unpredictable perf…

Progressive Cramming: Reliable Token Compression and What It Reveals

2026-07-23 · Dmitrii Tarasov, Timofei Lashukov, Elizaveta Goncharova, Andrey Kuznetsov arxiv

Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99\% accuracy thresholds leave it unclear whether residual errors reflect optimization failures or…

Sentence Compression as Deletion with Contextual Embeddings

2020-06-05 · Minh-Tien Nguyen, Bui Cong Minh, Dung Tien Le, Le Thai Linh

Sentence compression is the task of creating a shorter version of an input sentence while keeping important information. In this paper, we extend the task of compression by deletion with the use of contextual embeddings.…

SentenceSentence Compression

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

2024-11-06 · Charles Zhang, Benji Peng, Xintian Sun, Qian Niu 외

Word embeddings and language models have transformed natural language processing (NLP) by facilitating the representation of linguistic elements in continuous vector spaces. This review visits foundational concepts such …

Model CompressionSentenceTopic ModelsWord Embeddings

Light4GS: Lightweight Compact 4D Gaussian Splatting Generation via Context Model

2025-03-18 · Mufan Liu, Qi Yang, He Huang, Wenjie Huang 외

3D Gaussian Splatting (3DGS) has emerged as an efficient and high-fidelity paradigm for novel view synthesis. To adapt 3DGS for dynamic content, deformable 3DGS incorporates temporally deformable primitives with learnabl…

3DGSNovel View Synthesis