InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have made significant advancements in reasoning capabilities. However, they still face challenges such as high computational demands and privacy concerns. This paper focuses on developing efficient Small Language Models (SLMs) and Multimodal Small Language Models (MSLMs) that retain competitive reasoning abilities. We introduce a novel training pipeline that enhances reasoning capabilities and facilitates deployment on edge devices, achieving state-of-the-art performance while minimizing development costs. \InfR~ aims to advance AI systems by improving reasoning, reducing adoption barriers, and addressing privacy concerns through smaller model sizes. Resources are available at https://github. com/Reallm-Labs/InfiR.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
SwinFIR: Revisiting the SwinIR with Fast Fourier Convolution and Improved Training for Image Super-Resolution
Transformer-based methods have achieved impressive image restoration performance due to their capacities to model long-range dependency compared to CNN-based methods. However, advances like SwinIR adopts the window-based…
Data AugmentationImage ReconstructionImage RestorationImage Super-Resolution+2POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models
Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems de…
Multimodal ReasoningPrompt EngineeringUne plateforme de recommandation automatique d'emojis (An emoji recommandation platform)
Nous pr{\'e}sentons une interface de recommandation d{'}emojis porteurs de sentiments qui utilise un mod{\`e}le de pr{\'e}diction appris sur des messages informels priv{\'e}s. Chacun {\'e}tant associ{\'e} {\`a} deux scor…
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are…
Context-Informed Machine Translation of Manga using Multimodal Large Language Models
Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding …
Machine TranslationTranslation