paper-with-me

홈 › Papers

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

2024-12-19 · Jie Huang, Ruibing Hou, Jiahe Zhao, Hong Chang, Shiguang Shan

Human-centric perceptions play a crucial role in real-world applications. While recent human-centric works have achieved impressive progress, these efforts are often constrained to the visual domain and lack interaction with human instructions, limiting their applicability in broader scenarios such as chatbots and sports analysis. This paper introduces Referring Human Perceptions, where a referring prompt specifies the person of interest in an image. To tackle the new task, we propose RefHCM (Referring Human-Centric Model), a unified framework to integrate a wide range of human-centric referring tasks. Specifically, RefHCM employs sequence mergers to convert raw multimodal data -- including images, text, coordinates, and parsing maps -- into semantic tokens. This standardized representation enables RefHCM to reformulate diverse human-centric referring tasks into a sequence-to-sequence paradigm, solved using a plain encoder-decoder transformer architecture. Benefiting from a unified learning strategy, RefHCM effectively facilitates knowledge transfer across tasks and exhibits unforeseen capabilities in handling complex reasoning. This work represents the first attempt to address referring human perceptions with a general-purpose framework, while simultaneously establishing a corresponding benchmark that sets new standards for the field. Extensive experiments showcase RefHCM's competitive and even superior performance across multiple human-centric referring tasks. The code and data are publicly at https://github.com/JJJYmmm/RefHCM.

📄 PDF Abstract BibTeX arXiv:2412.14643

Code (1)

jjjymmm/refhcm 공식 구현 pytorch

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

UniHCP: A Unified Model for Human-Centric Perceptions

2023-03-06 · CVPR 2023 1 · Yuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang 외

Human-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have …

2D Pose EstimationAttributeHuman ParsingHuman Part Segmentation+7

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

2025-12-30 · Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 외 arxiv

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for u…

Referring Video Object Segmentation

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

2026-07-02 · Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai 외 arxiv

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…

Referring Expression

City landscape in sight: A crowdsourced framework for unlocking urban-scale window view perceptions from real estate imagery

2026-06-13 · Chucai Peng, Sijie Yang, Ang Liu, Yang Xiang 외 arxiv

City landscapes viewed through home windows influence quality of life, yet perceptions of actual window views at the urban scale remain understudied. This study presents an approach for large-scale mapping of perceptions…

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

2024-03-05 · CVPR 2024 1 · Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu 외

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabi…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2