paper-with-me

홈 › Papers

Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs

2023-10-01 · CVPR 2024 1 · Shiyu Xuan, Qingpei Guo, Ming Yang, Shiliang Zhang

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless, their performance in fine-grained image understanding tasks is still limited. To address this issue, this paper proposes a new framework to enhance the fine-grained image understanding abilities of MLLMs. Specifically, we present a new method for constructing the instruction tuning dataset at a low cost by leveraging annotations in existing datasets. A self-consistent bootstrapping method is also introduced to extend existing dense object annotations into high-quality referring-expression-bounding-box pairs. These methods enable the generation of high-quality instruction data which includes a wide range of fundamental abilities essential for fine-grained image perception. Moreover, we argue that the visual encoder should be tuned during instruction tuning to mitigate the gap between full image perception and fine-grained image perception. Experimental results demonstrate the superior performance of our method. For instance, our model exhibits a 5.2% accuracy improvement over Qwen-VL on GQA and surpasses the accuracy of Kosmos-2 by 24.7% on RefCOCO_val. We have also attained the top rank on the leaderboard of MMBench. This promising performance is achieved by training on only publicly available data, making it easily reproducible. The models, datasets, and codes are publicly available at https://github.com/SY-Xuan/Pink.

📄 PDF Abstract BibTeX arXiv:2310.00582

Code (1)

sy-xuan/pink 공식 구현 pytorch

Tasks

Referring Expression

Similar Papers 제목 키워드 기반

BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding

2026-01-04 · Hongbing Li, Linhui Xiao, Zihan Zhao, Qi Shen 외 arxiv

Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced …

Computational EfficiencyVisual Grounding

A simple model for pink noise from amplitude modulations

2023-01-26 · Masahiro Morikawa, Akika Nakamichi

We propose a simple model for the origin of pink noise (or 1/f fluctuation) based on the beat of cooperative waves. These cooperative waves arise spontaneously in a system with synchronization, resonance, and infrared di…

Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning

2019-08-16 · IJCNLP 2019 11 · Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith 외

Machine comprehension of texts longer than a single sentence often requires coreference resolution. However, most current reading comprehension benchmarks do not contain complex coreferential phenomena and hence fail to …

coreference-resolutionCoreference ResolutionReading ComprehensionSentence

Suppressing Pink Elephants with Direct Principle Feedback

2024-02-12 · Louis Castricato, Nathan Lile, Suraj Anand, Hailey Schoelkopf 외

Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable…

Language ModelingLanguage Modelling

The impact of physiological stress conditions on protein structure and trypsin inhibition of serine protease inhibitor Kazal type 1 (SPINK1) and its N34S variant

2020-12-18 · Ina Buchholz, Felix Nagel, Annelie Klein, Preshit R. Wagh 외

One of the most common mutations in the serine protease inhibitor Kazal type 1 (SPINK1) gene is the N34S variant which is strongly associated with chronic pancreatitis. Although it is assumed that N34S mutation constitut…