paper-with-me

Papers

Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI

2025-02-24 · Syed Abdul Gaffar Shakhadri, Kruthika KR, Kartik Basavaraj Angadi

We introduce Shakti VLM, a family of vision-language models in the capacity of 1B and 4B parameters designed to address data efficiency challenges in multimodal learning. While recent VLMs achieve strong performance through extensive training data, Shakti models leverage architectural innovations to attain competitive results with fewer tokens. Key advancements include QK-Normalization for attention stability, hybrid normalization techniques, and enhanced positional encoding. A three-stage training strategy further optimizes learning efficiency. Evaluations show that Shakti-Shakti-VLM-1B and Shakti-VLM-4B excel in document understanding, Visual Reasoning, OCR extraction, and general multimodal reasoning. Our results highlight that high performance can be achieved through model design and training strategy rather than sheer data volume, making Shakti an efficient solution for enterprise-scale multimodal tasks.

📄 PDF Abstract BibTeX arXiv:2502.17092

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingMultimodal ReasoningOptical Character Recognition (OCR)Visual Reasoning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments

2024-10-15 · Syed Abdul Gaffar Shakhadri, Kruthika KR, Rakshit Aralimatti

We introduce Shakti, a 2.5 billion parameter language model specifically optimized for resource-constrained environments such as edge devices, including smartphones, wearables, and IoT systems. Shakti combines high-perfo…

Language ModelingLanguage ModellingQuestion AnsweringSmall Language Model

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

2025-09-03 · Srihari Bandraupalli, Anupam Purwar arxiv

Open-source Vision-Language Models show immense promise for enterprise applications, yet a critical disconnect exists between academic evaluation and enterprise deployment requirements. Current benchmarks rely heavily on…

Object Detection

H2OVL-Mississippi Vision Language Models Technical Report

2024-10-17 · Shaikat Galib, Shanshan Wang, Guanshuo Xu, Pascal Pfeiffer 외

Smaller vision-language models (VLMs) are becoming increasingly important for privacy-focused, on-device applications due to their ability to run efficiently on consumer hardware for processing enterprise commercial docu…

Document AIVisual Question Answering

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning

2025-02-03 · Udita Ghosh, Dripta S. Raychaudhuri, Jiachen Li, Konstantinos Karydis 외

Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce PrefVLM, a framewor…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Transfer Learning

VLM3: Vision Language Models Are Native 3D Learners

2026-05-28 · Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu 외 arxiv

Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on exp…

Camera Pose EstimationDepth Estimation