paper-with-me

홈 › Papers

Parallel Context Windows for Large Language Models

2022-12-21 · Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham

When applied to processing long text, Large Language Models (LLMs) are limited by their context window. Existing efforts to address this limitation involve training specialized architectures, and cannot be easily applied to off-the-shelf LLMs. We present Parallel Context Windows (PCW), a method that alleviates the context window restriction for any off-the-shelf LLM without further training. The key to the approach is to carve a long context into chunks (``windows''), restrict the attention mechanism to apply only within each window, and re-use the positional embeddings across the windows. Our main results test the PCW approach on in-context learning with models that range in size between 750 million and 178 billion parameters, and show substantial improvements for tasks with diverse input and output spaces. We show additional benefits in other settings where long context windows may be beneficial: multi-hop questions and retrieval-augmented question answering with multiple retrieved documents. Our results highlight Parallel Context Windows as a promising method for applying off-the-shelf LLMs in a range of settings that require long text sequences. We make our code publicly available at https://github.com/ai21labs/parallel-context-windows.

📄 PDF Abstract BibTeX arXiv:2212.10947

Code (1)

AI21Labs/Parallel-Context-Windows 공식 구현 pytorch

Tasks

In-Context LearningPlaying the Game of 2048Question AnsweringRetrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Revisiting Parallel Context Windows: A Frustratingly Simple Alternative and Chain-of-Thought Deterioration

2023-05-24 · Kejuan Yang, Xiao Liu, Kaiwen Men, Aohan Zeng 외

We identify two crucial limitations in the evaluation of recent parallel-integrated method Parallel Context Windows (PCW), which extends the maximum context lengths of language models, e.g., 2048 for LLaMA, by harnessing…

Long-Context Understanding

Dehallucinating Parallel Context Extension for Retrieval-Augmented Generation

2024-12-19 · Zexiong Ma, Shengnan An, Zeqi Lin, Yanzhen Zou 외

Large language models (LLMs) are susceptible to generating hallucinated information, despite the integration of retrieval-augmented generation (RAG). Parallel context extension (PCE) is a line of research attempting to e…

HallucinationRAGRetrievalRetrieval-augmented Generation

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation

2026-02-04 · Luis Frentzen Salim, Esteban Carlin, Alexandre Morinvil, Xi Ai 외 arxiv

Building machine translation (MT) systems for low-resource languages is notably difficult due to the scarcity of high-quality data. Although Large Language Models (LLMs) have improved MT system performance, adapting them…

Machine Translation

Learning Adaptive Parallel Reasoning with Language Models

2025-04-21 · Jiayi Pan, Xiuyu Li, Long Lian, Charlie Snell 외

Scaling inference-time computation has substantially improved the reasoning capabilities of language models. However, existing methods have significant limitations: serialized chain-of-thought approaches generate overly …

4k

Pralekha: An Indic Document Alignment Evaluation Benchmark

2024-11-28 · Sanjay Suryanarayanan, Haiyue Song, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan 외

Mining parallel document pairs poses a significant challenge because existing sentence embedding models often have limited context windows, preventing them from effectively capturing document-level information. Another o…

SentenceSentence EmbeddingSentence-Embedding