GOO: A Dataset for Gaze Object Prediction in Retail Environments
One of the most fundamental and information-laden actions humans do is to look at objects. However, a survey of current works reveals that existing gaze-related datasets annotate only the pixel being looked at, and not the boundaries of a specific object of interest. This lack of object annotation presents an opportunity for further advancing gaze estimation research. To this end, we present a challenging new task called gaze object prediction, where the goal is to predict a bounding box for a person's gazed-at object. To train and evaluate gaze networks on this task, we present the Gaze On Objects (GOO) dataset. GOO is composed of a large set of synthetic images (GOO Synth) supplemented by a smaller subset of real images (GOO-Real) of people looking at objects in a retail environment. Our work establishes extensive baselines on GOO by re-implementing and evaluating selected state-of-the art models on the task of gaze following and domain adaptation. Code is available on github.
Code (1)
Tasks
Domain AdaptationGaze EstimationObjectSimilar Papers 제목 키워드 기반
TransGOP: Transformer-Based Gaze Object Prediction
Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, w…
Gaze EstimationObjectobject-detectionObject Detection+1Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenes
Recent advances in human motion prediction (HMP) have shifted focus from isolated motion data to integrating human-scene correlations. In particular, the latest methods leverage human gaze points, using their spatial…
Human motion predictionHuman-Object Interaction Detectionmotion predictionDriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
Predicting driver attention is a critical problem for developing explainable autonomous driving systems and understanding driver behavior in mixed human-autonomous vehicle traffic scenarios. Although significant progress…
Semantic SegmentationAutonomous DrivingPose2Gaze: Eye-body Coordination during Daily Activities for Gaze Prediction from Full-body Poses
Human eye gaze plays a significant role in many virtual and augmented reality (VR/AR) applications, such as gaze-contingent rendering, gaze-based interaction, or eye-based activity recognition. However, prior works on ga…
Activity RecognitionGaze PredictionHuman-Object Interaction DetectionGaze Estimation for Assisted Living Environments
Effective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated appro…
Gaze EstimationPose EstimationSemantic Segmentation