paper-with-me

홈 › Papers

ViRAC: A Vision-Reasoning Agent Head Movement Control Framework in Arbitrary Virtual Environments

2025-02-14 · Juyeong Hwang, Seong-Eun Hong, Hyeongyeop Kang

Creating lifelike virtual agents capable of interacting with their environments is a longstanding goal in computer graphics. This paper addresses the challenge of generating natural head rotations, a critical aspect of believable agent behavior for visual information gathering and dynamic responses to environmental cues. Although earlier methods have made significant strides, many rely on data-driven or saliency-based approaches, which often underperform in diverse settings and fail to capture deeper cognitive factors such as risk assessment, information seeking, and contextual prioritization. Consequently, generated behaviors can appear rigid or overlook critical scene elements, thereby diminishing the sense of realism. In this paper, we propose \textbf{ViRAC}, a \textbf{Vi}sion-\textbf{R}easoning \textbf{A}gent Head Movement \textbf{C}ontrol framework, which exploits the common-sense knowledge and reasoning capabilities of large-scale models, including Vision-Language Models (VLMs) and Large-Language Models (LLMs). Rather than explicitly modeling every cognitive mechanism, ViRAC leverages the biases and patterns internalized by these models from extensive training, thus emulating human-like perceptual processes without hand-tuned heuristics. Experimental results in multiple scenarios reveal that ViRAC produces more natural and context-aware head rotations than recent state-of-the-art techniques. Quantitative evaluations show a closer alignment with real human head-movement data, while user studies confirm improved realism and cognitive plausibility.

📄 PDF Abstract BibTeX arXiv:2502.10046

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense Reasoning

Similar Papers 제목 키워드 기반

DIJIT: A Robotic Head for an Active Observer

2025-12-08 · Mostafa Kamali Tabrizi, Mingshi Chi, Bir Bikram Dey, Kelly Yuan 외 arxiv

We present DIJIT, a novel binocular robotic head expressly designed for mobile agents that behave as active observers. DIJIT's unique breadth of functionality enables active vision research and the study of human-like ey…

Anticipation through Head Pose Estimation: a preliminary study

2024-08-10 · Federico Figari Tomenotti, Nicoletta Noceti

The ability to anticipate others' goals and intentions is at the basis of human-human social interaction. Such ability, largely based on non-verbal communication, is also a key to having natural and pleasant interactions…

Head Pose EstimationPose Estimation

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

2026-05-21 · Wenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li 외 arxiv

Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Languag…

Vision-Language Navigation

QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight

2026-04-03 · Ao Zhuang, Feng Yu, Tianbao Zhang, Linzuo Zhang 외 arxiv

We present QuadAgent, a training-free agent system for agile quadrotor flight guided by vision-language inputs. Unlike prior end-to-end or serial agent approaches, QuadAgent decouples high-level reasoning from low-level …

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

2025-08-11 · Tianyi Ma, Yue Zhang, Zehao Wang, Parisa Kordjamshidi arxiv

Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3D environments. While recent progress has been driven by large-scale pre-train…

Data Augmentation