paper-with-me

Papers

PaLI: A Jointly-Scaled Multilingual Language-Image Model

2022-09-14 · Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision, language, and multimodal tasks, in many languages. To train PaLI, we make use of large pre-trained encoder-decoder language models and Vision Transformers (ViTs). This allows us to capitalize on their existing capabilities and leverage the substantial cost of training them. We find that joint scaling of the vision and language components is important. Since existing Transformers for language are much larger than their vision counterparts, we train a large, 4-billion parameter ViT (ViT-e) to quantify the benefits from even larger-capacity vision models. To train PaLI, we create a large multilingual mix of pretraining tasks, based on a new image-text training set containing 10B images and texts in over 100 languages. PaLI achieves state-of-the-art in multiple vision and language tasks (such as captioning, visual question-answering, scene-text understanding), while retaining a simple, modular, and scalable design.

📄 PDF Abstract BibTeX arXiv:2209.06794

Code (1)

google-research/big_vision 공식 구현 jax

Tasks

DecoderFew-Shot Image ClassificationImage CaptioningImage ClassificationmodelQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningZero-Shot Image ClassificationZero-Shot Transfer Image Classification

Similar Papers 제목 키워드 기반

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

2023-10-13 · Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov 외

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this strong performance, we compare Vision Tra…

Chart Question AnsweringCross-Modal Retrievalimage-classificationImage Classification+6

PaLI-X: On Scaling up a Multilingual Vision and Language Model

2023-05-29 · Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa 외

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new leve…

Chart Question Answeringdocument understandingFine-Grained Image RecognitionIn-Context Learning+10

Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers

2024-09-29 · Prakash Dhakal, Daya Sagar Baral

Automatic text summarization in Nepali language is an unexplored area in natural language processing (NLP). Although considerable research has been dedicated to extractive summarization, the area of abstractive summariza…

Abstractive Text SummarizationArticlesExtractive SummarizationHeadline Generation+2

Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language

2026-02-27 · Nischal Karki, Bipesh Subedi, Prakash Poudyal, Rupak Raj Ghimire 외 arxiv

Transformer-based models such as BERT have significantly advanced Natural Language Processing (NLP) across many languages. However, Nepali, a low-resource language written in Devanagari script, remains relatively underex…

Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech

2026-06-21 · Samip Neupane, Sandesh Pokhrel, Sandesh Pyakurel, Basanta Joshi arxiv

Speaker diarization, the task of determining "who spoke when" in a multi-speaker recording, is a critical component in applications such as meeting transcription, accessibility tools, and multilingual information retriev…

Information RetrievalSpeaker Diarization