paper-with-me

홈 › Papers

Airavata: Introducing Hindi Instruction-tuned LLM

2024-01-26 · Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain, Aswanth Kumar M, Mohammed Safi Ur Rahman Khan, Diptesh Kanojia, Ratish Puduppully, Mitesh M. Khapra, Raj Dabre, Rudra Murthy, Anoop Kunchukuttan

We announce the initial release of "Airavata," an instruction-tuned LLM for Hindi. Airavata was created by fine-tuning OpenHathi with diverse, instruction-tuning Hindi datasets to make it better suited for assistive tasks. Along with the model, we also share the IndicInstruct dataset, which is a collection of diverse instruction-tuning datasets to enable further research for Indic LLMs. Additionally, we present evaluation benchmarks and a framework for assessing LLM performance across tasks in Hindi. Currently, Airavata supports Hindi, but we plan to expand this to all 22 scheduled Indic languages. You can access all artifacts at https://ai4bharat.github.io/airavata.

📄 PDF Abstract BibTeX arXiv:2401.15006

Code (1)

ai4bharat/indicinstruct 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis

2025-08-27 · Anusha Kamath, Kanishk Singla, Rakesh Paul, Raviraj Joshi 외 arxiv

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuanc…

Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs

2025-07-22 · Mahika Phutane, Aditya Vashistha arxiv

People with disabilities (PwD) experience disproportionately high levels of discrimination and hate online, particularly in India, where entrenched stigma and limited resources intensify these challenges. Large language …

Improving Multilingual Capabilities with Cultural and Local Knowledge in Large Language Models While Enhancing Native Performance

2025-04-13 · Ram Mohan Rao Kadiyala, Siddartha Pullakhandam, Siddhant Gupta, Drishti Sharma 외

Large Language Models (LLMs) have shown remarkable capabilities, but their development has primarily focused on English and other high-resource languages, leaving many languages underserved. We present our latest Hindi-E…

Minimal-Edit Instruction Tuning for Low-Resource Indic GEC

2025-11-28 · Akhil Rajeev P arxiv

Grammatical error correction for Indic languages faces limited supervision, diverse scripts, and rich morphology. We propose an augmentation-free setup that uses instruction-tuned large language models and conservative d…

parameter-efficient fine-tuningGrammatical Error Correction

PALO: A Polyglot Large Multimodal Model for 5B People

2024-02-22 · Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker, Salman Khan 외

In pursuit of more inclusive Vision-Language Models (VLMs), this study introduces a Large Multilingual Multimodal Model called PALO. PALO offers visual reasoning capabilities in 10 major languages, including English, Chi…

Language ModelingLanguage ModellingLarge Language ModelVisual Reasoning