paper-with-me

홈 › Papers

Surpassing Scale by Efficiency: A Compact 135M Parameter Foundational LLM Natively Adapted for the Bangla Language

2026-06-15 · Rabindra Nath Nandi arxiv

While the NLP landscape is dominated by multi-billion parameter architectures, their deployment in low-resource, non-Latin scripts remains computationally prohibitive for edge configurations, mobile systems, and decentralized local hardware. This paper presents bangla-smollm-135m, a highly compact 135-million parameter decoder-only foundational model engineered explicitly for high-efficiency language modeling in the Bangla script. By leveraging a deterministic intersect-and-append token merging strategy between TituLLMs and SmolLM2-135M, the model overcomes subword script fragmentation without destabilizing early pretrained parameter states. In zero-shot multi-task benchmark evaluations (PIQA_bn, OpenBookQA_bn, CommonsenseQA_bn, and Bangla_MMLU), bangla-smollm-135m matches or outperforms models twice its size (Gemma-3-270m) and achieves parity with models in the 1B parameter tier. The model is available at rnnandi/bangla-smollm-135m

📄 PDF Abstract BibTeX arXiv:2606.16383

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers

2025-07-03 · Pilsung Kang arxiv

Parameterized quantum circuits (PQCs) have recently emerged as promising components for enhancing the expressibility of neural architectures. In this work, we introduce QFFN-BERT, a hybrid quantum-classical transformer w…

Few-Shot Learning

CoSMoEs: Compact Sparse Mixture of Experts

2025-02-28 · Patrick Huber, Akshat Shrivastava, Ernie Chang, Chinnadhurai Sankar 외

Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device…

Mixture-of-Experts

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

2026-05-20 · Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen 외 arxiv

We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters across various benchmarks, while requiring s…

UltraFast-LiNET: Light-weight multi-scale shift convolutional network for real-time low-light image enhancement

2025-12-02 · Yuhan Chen, Yicui Shi, Guofa Li, Guangrui Bai 외 arxiv

Addressing the urgent need for high-performance real-time low-light image enhancement on resource-constrained edge devices in low-illumination scenarios such as nighttime and tunnels, this paper presents UltraFast-LiNET,…

Low-Light Image Enhancement

DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation

2025-03-13 · Chen Chen, Rui Qian, Wenze Hu, Tsu-Jui Fu 외

In this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protocols. We evaluate a range of DiT-based arc…

Image GenerationText to Image GenerationText-to-Image Generation