paper-with-me

홈 › Papers

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

2024-03-14 · Yufei Zhan, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang

Large Vision Language Models have achieved fine-grained object perception, but the limitation of image resolution remains a significant obstacle to surpass the performance of task-specific experts in complex and dense scenarios. Such limitation further restricts the model's potential to achieve nuanced visual and language referring in domains such as GUI Agents, Counting and \etc. To address this issue, we introduce a unified high-resolution generalist model, Griffon v2, enabling flexible object referring with visual and textual prompts. To efficiently scaling up image resolution, we design a simple and lightweight down-sampling projector to overcome the input tokens constraint in Large Language Models. This design inherently preserves the complete contexts and fine details, and significantly improves multimodal perception ability especially for small objects. Building upon this, we further equip the model with visual-language co-referring capabilities through a plug-and-play visual tokenizer. It enables user-friendly interaction with flexible target images, free-form texts and even coordinates. Experiments demonstrate that Griffon v2 can localize any objects of interest with visual and textual referring, achieve state-of-the-art performance on REC, phrase grounding, and REG tasks, and outperform expert models in object detection and object counting. Data, codes and models will be released at https://github.com/jefferyZhan/Griffon.

📄 PDF Abstract BibTeX arXiv:2403.09333

Code (1)

jefferyzhan/griffon 공식 구현 pytorch

Tasks

ObjectObject Countingobject-detectionObject DetectionPhrase Grounding

Similar Papers 제목 키워드 기반

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

2025-05-27 · Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng 외

Large Multimodal Models (LMMs) have recently demonstrated remarkable visual understanding performance on both vision-language and vision-centric tasks. However, they often fall short in integrating advanced, task-specifi…

Question AnsweringVisual Reasoning

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

2024-10-21 · Yufei Zhan, Hongyin Zhao, Yousong Zhu, Fan Yang 외

Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models typically focus on either vision-centric…

Instruction Followingobject-detectionObject DetectionQuestion Answering+5

Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models

2023-11-24 · Yufei Zhan, Yousong Zhu, Zhiyang Chen, Fan Yang 외

Replicating the innate human ability to detect all objects based on free-form texts at any granularity remains a formidable challenge for Large Vision Language Models (LVLMs). Current LVLMs are predominantly constrained …

AllReferring ExpressionReferring Expression Comprehension

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

2025-10-13 · Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino arxiv

We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support thre…

DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer

2025-04-28 · Junpeng Jiang, Gangyi Hong, Miao Zhang, Hengtong Hu 외

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealin…

Video Generation