paper-with-me

홈 › Papers

Motif 2 12.7B technical report

2025-11-07 · Junghwan Lim, Sungmin Lee, Dongseok Kim, Taehyun Kim, Eunhwan Park, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Wai Ting Cheung, Dahye Choi, Jaeheui Her, Jaeyeon Huh, Hanbin Jung, Changjin Kang, Beomgyu Kim, Minjae Kim, Taewhan Kim, Youngrok Kim, Hyukjin Kweon, Haesol Lee, Kungyu Lee, Dongpin Oh, Yeongjae Park, Bokki Ryu, Dongjoo Weon arxiv

We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding and robust instruction generalization under constrained compute budgets, Motif-2-12.7B builds upon Motif-2.6B with the integration of Grouped Differential Attention (GDA), which improves representational efficiency by disentangling signal and noise-control attention pathways. The model is pre-trained on 5.5 trillion tokens spanning diverse linguistic, mathematical, scientific, and programming domains using a curriculum-driven data scheduler that gradually changes the data composition ratio. The training system leverages the MuonClip optimizer alongside custom high-performance kernels, including fused PolyNorm activations and the Parallel Muon algorithm, yielding significant throughput and memory efficiency gains in large-scale distributed environments. Post-training employs a three-stage supervised fine-tuning pipeline that successively enhances general instruction adherence, compositional understanding, and linguistic precision. Motif-2-12.7B demonstrates competitive performance across diverse benchmarks, showing that thoughtful architectural scaling and optimized training design can rival the capabilities of much larger models.

📄 PDF Abstract BibTeX arXiv:2511.07464

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Motif 2.6B Technical Report

2025-08-02 · Junghwan Lim, Sungmin Lee, Dongseok Kim, Eunhwan Park 외 arxiv

Recent advancements in Large Language Models (LLMs) have revolutionized artificial intelligence, yet developing an effective foundational LLM that balances high performance with computational efficiency remains challengi…

Computational Efficiency

Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5

2018-10-31 · Avanti Shrikumar, Katherine Tian, Žiga Avsec, Anna Shcherbina 외

TF-MoDISco (Transcription Factor Motif Discovery from Importance Scores) is an algorithm for identifying motifs from basepair-level importance scores computed on genomic sequence data. This technical note focuses on vers…

Motif 3: Technical Report

2026-08-10 · Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung 외 hf

We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per to…

Long-Context UnderstandingReinforcement LearningMathematical ReasoningInstruction Following

Robust Machine Learning for Regulatory Sequence Modeling under Biological and Technical Distribution Shifts

2026-01-21 · Yiyao Yang arxiv

Robust machine learning for regulatory genomics is studied under biologically and technically induced distribution shifts. Deep convolutional and attention based models achieve strong in distribution performance on DNA r…

Finding Trolls Under Bridges: Preliminary Work on a Motif Detector

2022-04-12 · W. Victor H. Yarlott, Armando Ochoa, Anurag Acharya, Laurel Bobrow 외

Motifs are distinctive recurring elements found in folklore that have significance as communicative devices in news, literature, press releases, and propaganda. Motifs concisely imply a large constellation of culturally-…