paper-with-me

Papers

Towards Developing a Multilingual and Code-Mixed Visual Question Answering System by Knowledge Distillation

2021-09-10 · Findings (EMNLP) 2021 11 · Humair Raj Khan, Deepak Gupta, Asif Ekbal

Pre-trained language-vision models have shown remarkable performance on the visual question answering (VQA) task. However, most pre-trained models are trained by only considering monolingual learning, especially the resource-rich language like English. Training such models for multilingual setups demand high computing resources and multilingual language-vision dataset which hinders their application in practice. To alleviate these challenges, we propose a knowledge distillation approach to extend an English language-vision model (teacher) into an equally effective multilingual and code-mixed model (student). Unlike the existing knowledge distillation methods, which only use the output from the last layer of the teacher network for distillation, our student model learns and imitates the teacher from multiple intermediate layers (language and vision encoders) with appropriately designed distillation objectives for incremental knowledge extraction. We also create the large-scale multilingual and code-mixed VQA dataset in eleven different language setups considering the multiple Indian and European languages. Experimental results and in-depth analysis show the effectiveness of the proposed VQA model over the pre-trained language-vision models on eleven diverse language setups.

📄 PDF Abstract BibTeX arXiv:2109.04653

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

A Unified Framework for Multilingual and Code-Mixed Visual Question Answering

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we propose an effective deep learning framework for multilingual and code- mixed visual question answering. The pro- posed model is capable of predicting answers from the questions in Hindi, English or Cod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Marathi-English Code-mixed Text Generation

2023-09-28 · Dhiraj Amin, Sharvari Govilkar, Sagar Kulkarni, Yash Shashikant Lalit 외

Code-mixing, the blending of linguistic elements from distinct languages to form meaningful sentences, is common in multilingual settings, yielding hybrid languages like Hinglish and Minglish. Marathi, India's third most…

Text Generation

RetrieveGPT: Merging Prompts and Mathematical Models for Enhanced Code-Mixed Information Retrieval

2024-11-07 · Aniket Deroy, Subhankar Maity

Code-mixing, the integration of lexical and grammatical elements from multiple languages within a single sentence, is a widespread linguistic phenomenon, particularly prevalent in multilingual societies. In India, social…

Information RetrievalRetrievalSentence

Curriculum Script Distillation for Multilingual Visual Question Answering

2023-01-17 · Khyathi Raghavi Chandu, Alborz Geramifard

Pre-trained models with dual and cross encoders have shown remarkable success in propelling the landscape of several tasks in vision and language in Visual Question Answering (VQA). However, since they are limited by the…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Code Switched and Code Mixed Speech Recognition for Indic languages

2022-03-30 · Harveen Singh Chadha, Priyanshi Shah, Ankur Dhuriya, Neeraj Chhimwal 외

Training multilingual automatic speech recognition (ASR) systems is challenging because acoustic and lexical information is typically language specific. Training multilingual system for Indic languages is even more tough…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+1