paper-with-me

홈 › Papers

A Sovereign, Open-Source Foundation Model for German and English

2026-07-10 · The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering arxiv

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.

📄 PDF Abstract BibTeX arXiv:2607.09424

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sovereign AI: Rethinking Autonomy in the Age of Global Interdependence

2025-11-18 · Shalabh Kumar Singh, Shubhashis Sengupta arxiv

Artificial intelligence (AI) is emerging as a foundational general-purpose technology, raising new dilemmas of sovereignty in an interconnected world. While governments seek greater control over it, the very foundations …

Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian

2025-09-06 · Michael Hoffmann, Jophin John, Stefan Schweter, Gokul Ramakrishnan 외 arxiv

We present Llama-GENBA-10B, a trilingual foundation model addressing English-centric bias in large language models. Built on Llama 3.1-8B and scaled to 10B parameters, Llama-GENBA-10B is continuously pretrained on 164B t…

Cross-Lingual Transfer

Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs

2026-02-02 · Shaltiel Shmidman, Avi Shmidman, Amir DN Cohen, Moshe Koppel arxiv

Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training large language models (LLMs) for low-res…

SimpleNLG-DE: Adapting SimpleNLG 4 to German

2019-10-01 · WS 2019 10 · Daniel Braun, Kira Klimt, Daniela Schneider, Florian Matthes

SimpleNLG is a popular open source surface realiser for the English language. For German, however, the availability of open source and non-domain specific realisers is sparse, partly due to the complexity of the German l…

Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models

2026-01-26 · Kunat Pipatanakul, Pittawat Taveekitworachai arxiv

Large language models (LLMs) have progressed rapidly; however, most state-of-the-art models are trained and evaluated primarily in high-resource languages such as English and Chinese, and are often developed by a small n…

Legal Reasoning