paper-with-me

홈 › Papers

InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning

2025-02-17 · Congkai Xie, Shuo Cai, Wenjun Wang, Pengxiang Li, Zhijie Sang, Kejing Yang, Yiming Zhang, Zhen Li, Guanghao Zhu, Zeyu Liu, Yang Yu, Yuhang Liu, Su Lu, Baoyi He, Qi Zhou, Xiaotian Han, Jianbo Yuan, Shengyu Zhang, Fei Wu, Hongxia Yang

Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have made significant advancements in reasoning capabilities. However, they still face challenges such as high computational demands and privacy concerns. This paper focuses on developing efficient Small Language Models (SLMs) and Multimodal Small Language Models (MSLMs) that retain competitive reasoning abilities. We introduce a novel training pipeline that enhances reasoning capabilities and facilitates deployment on edge devices, achieving state-of-the-art performance while minimizing development costs. \InfR~ aims to advance AI systems by improving reasoning, reducing adoption barriers, and addressing privacy concerns through smaller model sizes. Resources are available at https://github. com/Reallm-Labs/InfiR.

📄 PDF Abstract BibTeX arXiv:2502.11573

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SwinFIR: Revisiting the SwinIR with Fast Fourier Convolution and Improved Training for Image Super-Resolution

2022-08-24 · Dafeng Zhang, Feiyu Huang, Shizhuo Liu, Xiaobing Wang 외

Transformer-based methods have achieved impressive image restoration performance due to their capacities to model long-range dependency compared to CNN-based methods. However, advances like SwinIR adopts the window-based…

Data AugmentationImage ReconstructionImage RestorationImage Super-Resolution+2

POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

2024-06-06 · Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu 외

Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems de…

Multimodal ReasoningPrompt Engineering

Une plateforme de recommandation automatique d'emojis (An emoji recommandation platform)

2017-06-01 · JEPTALNRECITAL 2017 6 · Ga{\"e}l Guibon, Magalie Ochs, Patrice Bellot

Nous pr{\'e}sentons une interface de recommandation d{'}emojis porteurs de sentiments qui utilise un mod{\`e}le de pr{\'e}diction appris sur des messages informels priv{\'e}s. Chacun {\'e}tant associ{\'e} {\`a} deux scor…

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

2025-10-11 · Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu 외 arxiv

Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are…

Context-Informed Machine Translation of Manga using Multimodal Large Language Models

2024-11-04 · Philip Lippmann, Konrad Skublicki, Joshua Tanner, Shonosuke Ishiwatari 외

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding …

Machine TranslationTranslation