paper-with-me

홈 › Papers

Chitrarth: Bridging Vision and Language for a Billion People

2025-02-21 · Shaharukh Khan, Ayush Tarun, Abhinav Ravi, Ali Faraz, Akshat Patidar, Praveen Kumar Pokala, Anagha Bhangare, Raja Kolla, Chandra Khatri, Shubham Agarwal

Recent multimodal foundation models are primarily trained on English or high resource European language data, which hinders their applicability to other medium and low-resource languages. To address this limitation, we introduce Chitrarth (Chitra: Image; Artha: Meaning), an inclusive Vision-Language Model (VLM), specifically targeting the rich linguistic diversity and visual reasoning across 10 prominent Indian languages. Our model effectively integrates a state-of-the-art (SOTA) multilingual Large Language Model (LLM) with a vision module, primarily trained on multilingual image-text data. Furthermore, we also introduce BharatBench, a comprehensive framework for evaluating VLMs across various Indian languages, ultimately contributing to more diverse and effective AI systems. Our model achieves SOTA results for benchmarks across low resource languages while retaining its efficiency in English. Through our research, we aim to set new benchmarks in multilingual-multimodal capabilities, offering substantial improvements over existing models and establishing a foundation to facilitate future advancements in this arena.

📄 PDF Abstract BibTeX arXiv:2502.15392

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityLanguage ModelingLanguage ModellingLarge Language ModelVisual Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Bridging Dictionary: AI-Generated Dictionary of Partisan Language Use

2024-07-12 · Hang Jiang, Doug Beeferman, William Brannon, Andrew Heyward 외

Words often carry different meanings for people from diverse backgrounds. Today's era of social polarization demands that we choose words carefully to prevent miscommunication, especially in political communication and j…

Language ModelingLanguage ModellingLarge Language Model

Innovative activities of Activision Blizzard: A patent network analysis

2025-02-04 · Artur F. Tomeczek

Microsoft's acquisition of Activision Blizzard valued at $68.7 billion has drastically altered the landscape of the video game industry. At the time of the takeover, the intellectual properties of Activision Blizzard inc…

Patent classificationStarcraft

Visual Clues: Bridging Vision and Language Foundations for Image Paragraph Captioning

2022-06-03 · Yujia Xie, Luowei Zhou, Xiyang Dai, Lu Yuan 외

People say, "A picture is worth a thousand words". Then how can we get the rich information out of the image? We argue that by using visual clues to bridge large pretrained vision foundation models and language models, w…

Image Paragraph CaptioningLanguage ModelingLanguage ModellingLarge Language Model

AWED-FiNER: Agents, Web applications, and Expert Detectors for Fine-grained Named Entity Recognition across 36 Languages for 6.6 Billion Speakers

2026-01-15 · Prachuryya Kaushik, Ashish Anand arxiv

Named Entity Recognition (NER) is a foundational task in Natural Language Processing (NLP) and Information Retrieval (IR), which facilitates semantic search and structured data extraction. We introduce \textbf{AWED-FiNER…

Information Retrieval

AI and Accessibility: A Discussion of Ethical Considerations

2019-08-21 · Meredith Ringel Morris

According to the World Health Organization, more than one billion people worldwide have disabilities. The field of disability studies defines disability through a social lens; people are disabled to the extent that socie…

speech-recognitionSpeech RecognitionTranslation