paper-with-me

Papers

Egocentric Hierarchical Visual Semantics

2023-05-09 · Luca Erculiani, Andrea Bontempelli, Andrea Passerini, Fausto Giunchiglia

We are interested in aligning how people think about objects and what machines perceive, meaning by this the fact that object recognition, as performed by a machine, should follow a process which resembles that followed by humans when thinking of an object associated with a certain concept. The ultimate goal is to build systems which can meaningfully interact with their users, describing what they perceive in the users' own terms. As from the field of Lexical Semantics, humans organize the meaning of words in hierarchies where the meaning of, e.g., a noun, is defined in terms of the meaning of a more general noun, its genus, and of one or more differentiating properties, its differentia. The main tenet of this paper is that object recognition should implement a hierarchical process which follows the hierarchical semantic structure used to define the meaning of words. We achieve this goal by implementing an algorithm which, for any object, recursively recognizes its visual genus and its visual differentia. In other words, the recognition of an object is decomposed in a sequence of steps where the locally relevant visual features are recognized. This paper presents the algorithm and a first evaluation.

📄 PDF Abstract BibTeX arXiv:2305.05422

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject Recognition

Similar Papers 제목 키워드 기반

ECO: Egocentric Cognitive Mapping

2018-12-02 · Jayant Sharma, Zixing Wang, Alberto Speranzon, Vijay Venkataraman 외

We present a new method to localize a camera within a previously unseen environment perceived from an egocentric point of view. Although this is, in general, an ill-posed problem, humans can effortlessly and efficiently …

Domain AdaptationNavigate

HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization

2025-08-30 · Joohyun Chang, Soyeon Hong, Hyogun Lee, Seong Jong Ha 외 arxiv

In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause sig…

Object LocalizationObject Recognition

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

2025-05-08 · Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng 외

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D vi…

3D visual groundingcross-modal alignmentVisual Grounding

EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

2024-05-22 · Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu 외

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semanti…

Human-Object Interaction DetectionObject

HCQA @ Ego4D EgoSchema Challenge 2024

2024-06-22 · Haoyu Zhang, Yuquan Xie, Yisen Feng, Zaijing Li 외

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comp…

Caption GenerationEgoSchemaMultiple-choice+2