paper-with-me

홈 › Papers

PaliGemma 2: A Family of Versatile VLMs for Transfer

2024-12-04 · Andreas Steiner, André Susano Pinto, Michael Tschannen, Daniel Keysers, Xiao Wang, Yonatan Bitton, Alexey Gritsenko, Matthias Minderer, Anthony Sherbondy, Shangbang Long, Siyang Qin, Reeve Ingle, Emanuele Bugliarello, Sahar Kazemzadeh, Thomas Mesnard, Ibrahim Alabdulmohsin, Lucas Beyer, Xiaohua Zhai

PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was also used by PaliGemma with the whole range of Gemma 2 models, from the 2B one all the way up to the 27B model. We train these models at three resolutions (224px, 448px, and 896px) in multiple stages to equip them with broad knowledge for transfer via fine-tuning. The resulting family of base models covering different model sizes and resolutions allows us to investigate factors impacting transfer performance (such as learning rate) and to analyze the interplay between the type of task, model size, and resolution. We further increase the number and breadth of transfer tasks beyond the scope of PaliGemma including different OCR-related tasks such as table structure recognition, molecular structure recognition, music score recognition, as well as long fine-grained captioning and radiography report generation, on which PaliGemma 2 obtains state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:2412.03555

Code (1)

kyutai-labs/moshivis jax

Tasks

Language ModelingLanguage ModellingOptical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

PaliGemma: A versatile 3B VLM for transfer

2024-07-10 · Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov 외

PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly knowledgeable base model that is effective t…

Language ModelingLanguage Modelling

DASH: Detection and Assessment of Systematic Hallucinations of VLMs

2025-03-30 · Maximilian Augustin, Yannic Neuhaus, Matthias Hein

Vision-language models (VLMs) are prone to object hallucinations, where they erroneously indicate the presenceof certain objects in an image. Existing benchmarks quantify hallucinations using relatively small, labeled da…

Object

Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma

2024-12-14 · Nouar AlDahoul, Myles Joshua Toledo Tan, Raghava Reddy Tera, Hezerul Abdul Karim 외

License plate recognition (LPR) involves automated systems that utilize cameras and computer vision to read vehicle license plates. Such plates collected through LPR can then be compared against databases to identify sto…

GPULicense Plate RecognitionOptical Character RecognitionOptical Character Recognition (OCR)

Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age

2024-10-31 · Nouar AlDahoul, Myles Joshua Toledo Tan, Harishwar Reddy Kasireddy, Yasir Zaki

Technologies for recognizing facial attributes like race, gender, age, and emotion have several applications, such as surveillance, advertising content, sentiment analysis, and the study of demographic trends and social …

AttributeEmotion ClassificationEmotion RecognitionSentiment Analysis

DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models

2025-08-25 · Abu Sufian, Anirudha Ghosh, Debaditya Barman, Marco Leo 외 arxiv

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities across various downstream tasks, including biometric face recognition (FR) with description. However, demographic biases remain a critical co…

Face Recognition