paper-with-me

홈 › Papers

Scaling Granite Code Models to 128K Context

2024-07-18 · Matt Stallone, Vaibhav Saxena, Leonid Karlinsky, Bridget McGinn, Tim Bula, Mayank Mishra, Adriana Meza Soria, Gaoyuan Zhang, Aditya Prasad, Yikang Shen, Saptha Surendran, Shanmukha Guttula, Hima Patel, Parameswaran Selvam, Xuan-Hong Dang, Yan Koyfman, Atin Sood, Rogerio Feris, Nirmit Desai, David D. Cox, Ruchir Puri, Rameswar Panda

This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code models from 2K/4K to 128K consists of a light-weight continual pretraining by gradually increasing its RoPE base frequency with repository-level file packing and length-upsampled long-context data. Additionally, we also release instruction-tuned models with long-context support which are derived by further finetuning the long context base models on a mix of permissively licensed short and long-context instruction-response pairs. While comparing to the original short-context Granite code models, our long-context models achieve significant improvements on long-context tasks without any noticeable performance degradation on regular code completion benchmarks (e.g., HumanEval). We release all our long-context Granite code models under an Apache 2.0 license for both research and commercial use.

📄 PDF Abstract BibTeX arXiv:2407.13739

Code (1)

ibm/data-prep-kit

Tasks

2k4kCode CompletionContinual PretrainingHumanEval

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Granite Guardian

2024-12-10 · Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury 외

We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with any large language model (LLM). These mo…

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+2

Granite Code Models: A Family of Open Foundation Models for Code Intelligence

2024-05-07 · Mayank Mishra, Matt Stallone, Gaoyuan Zhang, Yikang Shen 외

Large Language Models (LLMs) trained on code are revolutionizing the software development process. Increasingly, code LLMs are being integrated into software development environments to improve the productivity of human …

Code GenerationDecoder

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

2025-05-13 · George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano 외

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modality aligning the 2B and 8B parameter vari…

automatic-speech-translationBenchmarking

Granite Embedding R2 Models

2025-08-26 · Parul Awasthy, Aashka Trivedi, Yulong Li, Meet Doshi 외 arxiv

We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval applications. Building upon our first-genera…

Granite Embedding Models

2025-02-27 · Parul Awasthy, Aashka Trivedi, Yulong Li, Mihaela Bornea 외

We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, with both English and Multilingual capabili…

Information RetrievalKnowledge DistillationRetrieval