paper-with-me

홈 › Papers

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

2026-01-05 · YuanLab. ai, :, Shawn Wu, Sean Wang, Louie Li, Darcy Chen, Allen Wang, Jiangang Luo, Xudong Zhao, Joseph Shen, Gawain Ma, Jasper Jia, Marcus Mao, Claire Wang, Hunter He, Carol Wang, Zera Zhang, Jason Wang, Chonly Shen, Leo Zhang, Logan Chen, Qasim Meng, James Gong, Danied Zhao, Penn Zheng, Owen Zhu, Tong Yu arxiv

We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks while maintaining competitive capabilities on general-purpose tasks. To address the overthinking phenomenon commonly observed in Large Reasoning Models (LRMs), we propose Reflection-aware Adaptive Policy Optimization (RAPO), a novel RL training algorithm that effectively regulates overthinking behaviors. In enterprise-oriented tasks such as retrieval-augmented generation (RAG), complex table understanding, and summarization, Yuan3.0 Flash consistently achieves superior performance. Moreover, it also demonstrates strong reasoning capabilities in domains such as mathematics, science, etc., attaining accuracy comparable to frontier model while requiring only approximately 1/4 to 1/2 of the average tokens. Yuan3.0 Flash has been fully open-sourced to facilitate further research and real-world deployment: https://github.com/Yuan-lab-LLM/Yuan3.0.

📄 PDF Abstract BibTeX arXiv:2601.01718

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

2024-05-14 · Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong 외

We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully design the transformer structure, text encoder, and positi…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+2

HunyuanImage 3.0 Technical Report

2025-09-28 · Tencent Hunyuan Foundation Model Team arxiv

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of Huny…

Image Generation

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

2026-07-06 · Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang 외 arxiv

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document unde…

Information ExtractionText Spotting

LongCat-Flash-Omni Technical Report

2025-10-31 · Meituan LongCat Team, Bairui Wang, Bayan, Bin Xiao 외 arxiv

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strate…

FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression

2025-01-01 · CVPR 2025 1 · Bo Tong, Bokai Lai, Yiyi Zhou, Gen Luo 외

Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLM…

Descriptive