paper-with-me

Papers

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

2026-08-07 · Yun Li, Biao Yang, Peixi Wu, Yunhao Zhou, Mingzhou Jiang, Wei Yuan, Fan Yang, Wenwu Ou arxiv

Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess representations through discriminative tasks or geometric criteria centered on separability in embedding space. However, strong performance on such evaluations does not establish whether content compressed into an embedding remains accessible to a downstream generator. To address this gap, we introduce the Generative Embedding Benchmark (GEB), in which a decoder answers questions using only a frozen embedding and question text, without access to the original image or intermediate visual features. Answer quality under this readout measures generative information: the answer-relevant content recoverable from an embedding. GEB includes a curated visual-question-answering dataset with a 1,800-item development split and a held-out 900-item test split covering natural images, scene text, and visual documents. Using a common decoder and training recipe, we evaluate seven public embedding models in visual-only and vision-language joint modes. On the test set, visual-only scores range from 28.25 to 33.21; with image-question joint encoding, all five VLM-based embedding models score higher, and the best reaches 65.56. Matched embeddings also outperform text-only inputs, zero embeddings, and shuffled embeddings. Natural-image information is much easier to recover than scene text or visual-document information, while a Qwen3-VL-2B reference with access to the original image reaches 84.30. Together, these results show that generative readout exposes information bottlenecks that separability-based evaluation does not capture.

📄 PDF Abstract BibTeX arXiv:2608.06972

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PreFIQs: Face Image Quality Is What Survives Pruning

2026-05-13 · Jan Niklas Kolf, Guray Ozgur, Andrea Atzori, Žiga Babnik 외 arxiv

Face Image Quality Assessment (FIQA) evaluates the utility of a face image for automated face recognition (FR) systems. In this work, we propose PreFIQs, an unsupervised and training-free FIQA framework grounded in the P…

Face Image Quality AssessmentFace Recognition

Information Bottleneck Constrained Latent Bidirectional Embedding for Zero-Shot Learning

2020-09-16 · Yang Liu, Lei Zhou, Xiao Bai, Lin Gu 외

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Though many ZSL methods rely on a direct mapping between the visual and the semantic space,…

AttributeZero-Shot Learning

Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software

2026-06-26 · Daniel Russo arxiv

Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates them the way it has always evaluated components, one agent at a time, on isolated benchmark tasks. Yet ag…

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

2026-05-13 · Vardhan Dongre, Joseph Hsieh, Viet Dac Lai, Seunghyun Yoon 외 arxiv

Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instructions, persona, and rules. This degradation has been measured behaviorall…

Robust information propagation through noisy neural circuits

2017-02-26

Sensory neurons give highly variable responses to stimulation, which can limit the amount of stimulus information available to downstream circuits. Much work has investigated the factors that affect the amount of informa…

Informativeness