paper-with-me

홈 › Papers

Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP

2024-10-11 · Eunji Kim, Kyuhong Shim, Simyung Chang, Sungroh Yoon

A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual elements within a sentence depending on the context, efforts to account for variation of importance in constructing text embeddings have been lacking. We propose a framework of Semantic Token Reweighting to build Interpretable text embeddings (SToRI), which incorporates controllability as well. SToRI refines the text encoding process in CLIP by differentially weighting semantic elements based on contextual importance, enabling finer control over emphasis responsive to data-driven insights and user preferences. The efficacy of SToRI is demonstrated through comprehensive experiments on few-shot image classification and image retrieval tailored to user preferences.

📄 PDF Abstract BibTeX arXiv:2410.08469

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Image Classificationimage-classificationImage ClassificationImage RetrievalRetrievalSentence

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Semantic Fusion with Fuzzy-Membership Features for Controllable Language Modelling

2025-09-14 · Yongchao Huang, Hassan Raza arxiv

We propose semantic fusion, a lightweight scheme that augments a Transformer language model (LM) with a parallel, fuzzy-membership feature channel that encodes token-level semantics. Each token is represented by a vector…

Language Modelling

iBERT: Interpretable Embeddings via Sense Decomposition

2025-10-10 · Vishal Anand, Milad Alshomary, Kathleen McKeown arxiv

We present iBERT (interpretable-BERT), an encoder to produce inherently interpretable and controllable embeddings - designed to modularize and expose the discriminative cues present in language, such as semantic or styli…

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

2026-06-29 · Shuochen Chang, Qingyang Liu, Shaobo Wang, Bingjie Gao 외 arxiv

Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. Latent reasoning reduces this cost by shif…

Reinforcement Learning

Dynamic Eraser for Guided Concept Erasure in Diffusion Models

2026-04-13 · Qinghui Gong arxiv

Concept erasure in Text-To-Image (T2I) diffusion models is vital for safe content generation, but existing inference-time methods face significant limitations. Feature-correction approaches often cause uncontrolled over-…

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

2025-10-27 · Nathan Paek, Yongyi Zang, Qihui Yang, Randal Leistikow arxiv

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semant…

Audio GenerationMusic Generation