paper-with-me

홈 › Papers

Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings

2025-11-18 · Xueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins, Yozen Liu, Leman Akoglu, Neil Shah, Tong Zhao arxiv

Large language models produce powerful text embeddings, but their causal attention mechanism restricts the flow of information from later to earlier tokens, degrading representation quality. While recent methods attempt to solve this by prepending a single summary token, they over-compress information, hence harming performance on long documents. We propose Hierarchical Token Prepending (HTP), a method that resolves two critical bottlenecks. To mitigate attention-level compression, HTP partitions the input into blocks and prepends block-level summary tokens to subsequent blocks, creating multiple pathways for backward information flow. To address readout-level over-squashing, we replace last-token pooling with mean-pooling, a choice supported by theoretical analysis. HTP achieves consistent performance gains across 11 retrieval datasets and 30 general embedding benchmarks, especially in long-context settings. As a simple, architecture-agnostic method, HTP enhances both zero-shot and finetuned models, offering a scalable route to superior long-document embeddings.

📄 PDF Abstract BibTeX arXiv:2511.14868

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMs

2024-12-16 · Yuchen Fu, Zifeng Cheng, Zhiwei Jiang, Zhonghui Wang 외

Extracting sentence embeddings from large language models (LLMs) is a promising direction, as LLMs have demonstrated stronger semantic understanding capabilities. Previous studies typically focus on prompt engineering to…

Prompt EngineeringSemantic Textual SimilaritySentenceSentence Embedding+3

FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference

2024-05-07 · Runheng Liu, Xingchen Xiao, Heyan Huang, Zewen Chi 외

Retrieval-Augmented Language Modeling (RALM) by integrating large language models (LLM) with relevant documents from an external corpus is a proven method for enabling the LLM to generate information beyond the scope of …

Language ModelingLanguage ModellingRetrieval

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

2025-03-17 · Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen 외

The remarkable performance of large multimodal models (LMMs) has attracted significant interest from the image segmentation community. To align with the next-token-prediction paradigm, current LMM-driven segmentation met…

Image SegmentationSegmentationSemantic SegmentationVisual Grounding

HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models

2025-03-14 · Ziqin Zhou, Yifan Yang, Yuqing Yang, Tianyu He 외

Text-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain g…

Text-to-Video GenerationVideo Generation

FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

2025-02-16 · Hui Wang, Shujie Liu, Lingwei Meng, Jinyu Li 외

To advance continuous-valued token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive …

Language ModelingLanguage ModellingSpeech Synthesis