paper-with-me

홈 › Papers

Nemotron-4 15B Technical Report

2024-02-26 · Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, Vibhu Jawa, Jiwei Liu, Ameya Mahabaleshwarkar, Osvald Nitski, Annika Brundyn, James Maki, Miguel Martinez, Jiaxuan You, John Kamalu, Patrick Legresley, Denys Fridman, Jared Casper, Ashwath Aithal, Oleksii Kuchaiev, Mohammad Shoeybi, Jonathan Cohen, Bryan Catanzaro

We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assessed on English, multilingual, and coding tasks: it outperforms all existing similarly-sized open models on 4 out of 7 downstream evaluation areas and achieves competitive performance to the leading open models in the remaining ones. Specifically, Nemotron-4 15B exhibits the best multilingual capabilities of all similarly-sized models, even outperforming models over four times larger and those explicitly specialized for multilingual tasks.

📄 PDF Abstract BibTeX arXiv:2402.16819

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Nemotron-4 340B Technical Report

2024-06-17 · Nvidia, :, Bo Adler, Niket Agarwal 외

We release the Nemotron-4 340B model family, including Nemotron-4-340B-Base, Nemotron-4-340B-Instruct, and Nemotron-4-340B-Reward. Our models are open access under the NVIDIA Open Model License Agreement, a permissive mo…

Synthetic Data Generation

Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery

2026-01-27 · Meng Xin, Sweta Priyadarshi, Jingyu Xin, Bilal Kartal 외 arxiv

This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-p…

Reinforcement Learning

NVIDIA Nemotron 3: Efficient and Open Intelligence

2025-12-24 · NVIDIA, :, Aaron Blakeman, Aaron Grattafiori 외 arxiv

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transfo…

Reinforcement LearningText Generation

Jupiter-N Technical Report

2026-04-19 · George Drayson arxiv

We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three objectives: (1) agentic capability via uncertainty-curated trajectories; (…

Instruction Following

Llama-Nemotron: Efficient Reasoning Models

2025-05-02 · Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah 외

We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an open license for enterprise use. The family…

Knowledge DistillationNeural Architecture Search