paper-with-me

Papers

GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model

2025-01-01 · CVPR 2025 1 · Yue Han, Jiangning Zhang, Junwei Zhu, Runze Hou, Xiaozhong Ji, Chuming Lin, Xiaobin Hu, Zhucun Xue, Yong liu

Multimodal Language Learning Models (MLLMs) have shown remarkable performance in image understanding, generation, and editing, with recent advancements achieving pixel-level grounding with reasoning. However, these models for common objects struggle with fine-grained face understanding. In this work, we introduce the FacePlayGround-240K dataset, the first pioneering large-scale, pixel-grounded face caption and question-answer (QA) dataset, meticulously curated for alignment pretraining and instruction-tuning. We present the GroundingFace framework, specifically designed to enhance fine-grained face understanding. This framework significantly augments the capabilities of existing grounding models in face part segmentation, face attribute comprehension, while preserving general scene understanding. Comprehensive experiments validate that our approach surpasses current state-of-the-art models in pixel-grounded face captioning/QA and various downstream tasks, including face captioning, referring segmentation, and zero-shot face attribute recognition.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelScene Understanding

Similar Papers 제목 키워드 기반

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

2026-01-06 · Yankai Jiang, Qiaoru Li, Binlu Xu, Haoran Sun 외 arxiv

Recent research on medical MLLMs has gradually shifted its focus from image-level understanding to fine-grained, pixel-level comprehension. Although segmentation serves as the foundation for pixel-level understanding, ex…

Reinforcement LearningVisual Reasoning

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

2025-09-22 · Ye Liu, Zongyang Ma, Junfu Pu, Zhongang Qi 외 arxiv

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conver…

Referring Expression SegmentationQuestion AnsweringVisual Reasoning

PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation

2024-12-19 · Muntasir Wahed, Kiet A. Nguyen, Adheesh Sunil Juvekar, Xinzhuo Li 외

Despite significant advancements in Large Vision-Language Models (LVLMs), existing pixel-grounding models operate on single-image settings, limiting their ability to perform detailed, fine-grained comparisons across mult…

Reasoning Segmentation

SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories

2025-03-11 · CVPR 2025 1 · Muzhi Zhu, Yuzhuo Tian, Hao Chen, Chunluan Zhou 외

While MLLMs have demonstrated adequate image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks like VQA and visual grounding r…

Decision MakingInteractive SegmentationSegmentationVisual Grounding+2

SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories

2025-03-11 · Zhu, Muzhi; Tian, Yuzhuo; Chen, Hao; Zhou 외

While MLLMs have demonstrated adequate image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks like VQA and visual grounding r…

Decision MakingInteractive SegmentationReferring Expression SegmentationSegmentation+3