paper-with-me

홈 › Papers

Referring to Any Person

2025-03-11 · Qing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren, Yuda Xiong, Yihao Chen, Qin Liu, Lei Zhang

Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referring to any person, holds substantial practical value. However, we find that existing models generally fail to achieve real-world usability, and current benchmarks are limited by their focus on one-to-one referring, that hinder progress in this area. In this work, we revisit this task from three critical perspectives: task definition, dataset design, and model architecture. We first identify five aspects of referable entities and three distinctive characteristics of this task. Next, we introduce HumanRef, a novel dataset designed to tackle these challenges and better reflect real-world applications. From a model design perspective, we integrate a multimodal large language model with an object detection framework, constructing a robust referring model named RexSeek. Experimental results reveal that state-of-the-art models, which perform well on commonly used benchmarks like RefCOCO/+/g, struggle with HumanRef due to their inability to detect multiple individuals. In contrast, RexSeek not only excels in human referring but also generalizes effectively to common object referring, making it broadly applicable across various perception tasks. Code is available at https://github.com/IDEA-Research/RexSeek

📄 PDF Abstract BibTeX arXiv:2503.08507

Code (1)

idea-research/rexseek 공식 구현 pytorch

Tasks

Large Language ModelMultimodal Large Language Modelobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

2026-08-28 · Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan 외 arxiv

Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…

Referring Expression

RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D

2023-08-23 · ICCV 2023 1 · Shuhei Kurita, Naoki Katsura, Eri Onami

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capa…

ObjectObject TrackingReferring ExpressionReferring Expression Comprehension

Referring Atomic Video Action Recognition

2024-07-02 · Kunyu Peng, Jia Fu, Kailun Yang, Di Wen 외

We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task dif…

Action LocalizationAction RecognitionQuestion AnsweringReferring Expression+3

A case study on context-bound referring expression generation

2019-10-01 · WS 2019 10 · Maurice Langner

In recent years, Bayesian models of referring expression generation have gained prominence in order to produce situationally more adequate referring expressions. Basically, these models enable the integration of differen…

Referring ExpressionReferring expression generation

Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing

2019-03-03 · CVPR 2019 6 · Xihui Liu, ZiHao Wang, Jing Shao, Xiaogang Wang 외

Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types of information from visual and textual …

Referring Expression