paper-with-me

Papers

Towards Unconstrained Human-Object Interaction

2026-04-15 · Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci arxiv

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in-the-wild HOI detection. We define the Unconstrained HOI (U-HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test-time inference and language-to-graph conversion to extract structured interactions from free-form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U-HOI. Code will be available at https://github.com/francescotonini/anyhoi

📄 PDF Abstract BibTeX arXiv:2604.14069

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos

2025-09-20 · Haoyang Weng, Yitang Li, Nikhil Sobanbabu, Zihan Wang 외 arxiv

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framewo…

Reinforcement Learning

Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction

2018-06-11 · Mohit Shridhar, David Hsu

This paper presents INGRESS, a robot system that follows human natural language instructions to pick and place everyday objects. The core issue here is the grounding of referring expressions: infer objects and their rela…

Question GenerationQuestion-GenerationVisual Grounding

HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video

2023-11-30 · CVPR 2024 1 · Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Muhammed Kocabas 외

Since humans interact with diverse objects every day, the holistic 3D capture of these interactions is important to understand and model human behaviour. However, most existing methods for hand-object reconstruction from…

3D ReconstructionObjectObject Reconstruction

Next-Active-Object prediction from Egocentric Videos

2019-04-10 · Antonino Furnari, Sebastiano Battiato, Kristen Grauman, Giovanni Maria Farinella

Although First Person Vision systems can sense the environment from the user's perspective, they are generally unable to predict his intentions and goals. Since human activities can be decomposed in terms of atomic actio…

ObjectPrediction

Weakly-Supervised Physically Unconstrained Gaze Estimation

2021-05-20 · CVPR 2021 1 · Rakshit Kothari, Shalini De Mello, Umar Iqbal, Wonmin Byeon 외

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environ…

Domain GeneralizationGaze Estimation