Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
We introduce Shakti VLM, a family of vision-language models in the capacity of 1B and 4B parameters designed to address data efficiency challenges in multimodal learning. While recent VLMs achieve strong performance through extensive training data, Shakti models leverage architectural innovations to attain competitive results with fewer tokens. Key advancements include QK-Normalization for attention stability, hybrid normalization techniques, and enhanced positional encoding. A three-stage training strategy further optimizes learning efficiency. Evaluations show that Shakti-Shakti-VLM-1B and Shakti-VLM-4B excel in document understanding, Visual Reasoning, OCR extraction, and general multimodal reasoning. Our results highlight that high performance can be achieved through model design and training strategy rather than sheer data volume, making Shakti an efficient solution for enterprise-scale multimodal tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
document understandingMultimodal ReasoningOptical Character Recognition (OCR)Visual ReasoningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments
We introduce Shakti, a 2.5 billion parameter language model specifically optimized for resource-constrained environments such as edge devices, including smartphones, wearables, and IoT systems. Shakti combines high-perfo…
Language ModelingLanguage ModellingQuestion AnsweringSmall Language ModelVLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
Open-source Vision-Language Models show immense promise for enterprise applications, yet a critical disconnect exists between academic evaluation and enterprise deployment requirements. Current benchmarks rely heavily on…
Object DetectionH2OVL-Mississippi Vision Language Models Technical Report
Smaller vision-language models (VLMs) are becoming increasingly important for privacy-focused, on-device applications due to their ability to run efficiently on consumer hardware for processing enterprise commercial docu…
Document AIVisual Question AnsweringPreference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce PrefVLM, a framewor…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Transfer LearningVLM3: Vision Language Models Are Native 3D Learners
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on exp…
Camera Pose EstimationDepth Estimation