paper-with-me

Papers

Bielik 11B v2 Technical Report

2025-05-05 · Krzysztof Ociepa, Łukasz Flis, Krzysztof Wróbel, Adrian Gwoździej, Remigiusz Kinas

We present Bielik 11B v2, a state-of-the-art language model optimized for Polish text processing. Built on the Mistral 7B v0.2 architecture and scaled to 11B parameters using depth up-scaling, this model demonstrates exceptional performance across Polish language benchmarks while maintaining strong cross-lingual capabilities. We introduce two key technical innovations: Weighted Instruction Cross-Entropy Loss, which optimizes learning across diverse instruction types by assigning quality-based weights to training examples, and Adaptive Learning Rate, which dynamically adjusts based on context length. Comprehensive evaluation across multiple benchmarks demonstrates that Bielik 11B v2 outperforms many larger models, including those with 2-6 times more parameters, and significantly surpasses other specialized Polish language models on tasks ranging from linguistic understanding to complex reasoning. The model's parameter efficiency and extensive quantization options enable deployment across various hardware configurations, advancing Polish language AI capabilities and establishing new benchmarks for resource-efficient language modeling in less-represented languages.

📄 PDF Abstract BibTeX arXiv:2505.02410

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingQuantization

Similar Papers 제목 키워드 기반

Making Bielik LLM Reason (Better): A Field Report

2026-03-11 · Adam Trybus, Bartosz Bartnicki, Remigiusz Kinas arxiv

This paper presents a research program dedicated to evaluating and advancing the reasoning capabilities of Bielik, a Polish large language model. The study describes a number of stages of work: initial benchmarking and c…

Bielik v3 Small: Technical Report

2025-05-05 · Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel 외

We introduce Bielik v3, a series of parameter-efficient generative text models (1.5B and 4.5B) optimized for Polish language processing. These models demonstrate that smaller, well-optimized architectures can achieve per…

Language ModelingLanguage Modelling

Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language

2026-03-12 · Remigiusz Kinas, Paweł Kiszczak, Sergio P. Perez, Krzysztof Ociepa 외 arxiv

This report details the creation of Bielik-Minitron-7B, a compressed 7.35B parameter version of the Bielik-11B-v3.0 model, specifically optimized for European languages. By leveraging a two-stage compression methodology …

Knowledge DistillationReinforcement Learning

Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series

2026-04-12 · Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel 외 arxiv

The development of the Bielik v3 PL series, encompassing both the 7B and 11B parameter variants, represents a significant milestone in the field of language-specific large language model (LLM) optimization. While general…

Reinforcement Learning

Bielik 11B v3: Multilingual Large Language Model for European Languages

2025-12-30 · Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel 외 arxiv

We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 archite…

Reinforcement Learning