paper-with-me

홈 › Papers

Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors

2026-01-23 · Chen-Ying Chien, Po-Chih Kuo arxiv

Understanding how people perceive and evaluate interior spaces is essential for designing environments that promote well-being. However, predicting aesthetic experiences remains difficult due to the subjective nature of perception and the complexity of visual responses. This study introduces a dual-branch CNN-LSTM framework that fuses visual features with eye-tracking signals to predict aesthetic evaluations of residential interiors. We collected a dataset of 224 interior design videos paired with synchronized gaze data from 28 participants who rated 15 aesthetic dimensions. The proposed model attains 72.2% accuracy on objective dimensions (e.g., light) and 66.8% on subjective dimensions (e.g., relaxation), outperforming state-of-the-art video baselines and showing clear gains on subjective evaluation tasks. Notably, models trained with eye-tracking retain comparable performance when deployed with visual input alone. Ablation experiments further reveal that pupil responses contribute most to objective assessments, while the combination of gaze and visual cues enhances subjective evaluations. These findings highlight the value of incorporating eye-tracking as privileged information during training, enabling more practical tools for aesthetic assessment in interior design.

📄 PDF Abstract BibTeX arXiv:2601.16811

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition

2026-02-27 · Jae Young Choi, Seon Gyeom Kim, Hyungjun Yoon, Taeckyung Lee 외 arxiv

Large Language Models (LLMs) have emerged as foundation models for IoT applications such as human activity recognition (HAR). However, directly applying high-frequency and multi-dimensional sensor data, such as eye-track…

Human Activity Recognition

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

2025-12-10 · Yuyang Li, Yinghan Chen, Zihang Zhao, Puhao Li 외 arxiv

Robotic manipulation requires both rich multimodal perception and effective learning frameworks to handle complex real-world tasks. See-through-skin (STS) sensors, which combine tactile and visual perception, offer promi…

Robot ManipulationContact Detection

Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering Pairs

2023-10-26 · Yuxin Zuo, Bei Li, Chuanhao Lv, Tong Zheng 외

This paper presents an in-depth study of multimodal machine translation (MMT), examining the prevailing understanding that MMT systems exhibit decreased sensitivity to visual information when text inputs are complete. In…

AttributeMachine TranslationMultimodal Machine TranslationQuestion Answering+2

MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors

2024-03-26 · CVPR 2024 1 · He Zhang, Shenghao Ren, Haolei Yuan, Jianhui Zhao 외

Foot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. H…

Translation

Reliable Object Tracking by Multimodal Hybrid Feature Extraction and Transformer-Based Fusion

2024-05-28 · Hongze Sun, Rui Liu, Wuque Cai, Jun Wang 외

Visual object tracking, which is primarily based on visible light image sequences, encounters numerous challenges in complicated scenarios, such as low light conditions, high dynamic ranges, and background clutter. To ad…

ObjectObject TrackingVisual Object Tracking