paper-with-me

Papers

Arcee Trinity Large Technical Report

2026-02-19 · Varun Singh, Lucas Krauss, Sami Jaghouar, Matej Sirovatka, Charles Goddard, Fares Obied, Jack Min Ong, Jannik Straube, Fern, Aria Harley, Conner Stewart, Colin Kealty, Maziyar Panahi, Simon Kirsten, Anushka Deshpande, Anneketh Vij, Arthur Bresnu, Pranav Veldurthi, Raghav Ravishankar, Hardik Bishnoi, DatologyAI Team, Arcee AI Team, Prime Intellect Team, Mark McQuade, Johannes Hagemann, Lucas Atkins arxiv

We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai.

📄 PDF Abstract BibTeX arXiv:2602.17004

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models

2025-05-23 · Xuchen Pan, Yanxi Chen, Yushuo Chen, Yuchang Sun 외

Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models. It is built with a decoupled design, consisting of (1) an RFT-core that unifies and…

Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation

2024-06-21 · Shamane Siriwardhana, Mark McQuade, Thomas Gauthier, Lucas Atkins 외

We conducted extensive experiments on domain adaptation of the Meta-Llama-3-70B-Instruct model on SEC data, exploring its performance on both general and domain-specific benchmarks. Our focus included continual pre-train…

Domain AdaptationLanguage ModelingLanguage Modelling

TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems

2026-03-16 · Kai Wang, Biaojie Zeng, Zeming Wei, Chang Jin 외 arxiv

With the rapid development of LLM-based multi-agent systems (MAS), their significant safety and security concerns have emerged, which introduce novel risks going beyond single agents or LLMs. Despite attempts to address …

Trinity: A No-Code AI platform for complex spatial datasets

2021-06-21 · C. V. Krishnakumar Iyer, Feili Hou, Henry Wang, Yonghong Wang 외

We present a no-code Artificial Intelligence (AI) platform called Trinity with the main design goal of enabling both machine learning researchers and non-technical geospatial domain experts to experiment with domain-spec…

Feature EngineeringSemantic Segmentation

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

2024-08-06 · Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu 외

This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These mul…

Medical Visual Question AnsweringOrgan DetectionRetrieval-augmented GenerationVisual Question Answering (VQA)