paper-with-me

Papers

A Vision Based Framework Integrating Attention and Action Cues for Interpretable Cognitive Workload Assessment in Human-Robot Collaborative Assembly

2026-09-14 · Junyan Xiong, Naiyi Feng, Xingke Xia, Qihang Fan, Suchang Chen, Daqiang Guo arxiv

The introduction of human-robot collaboration (HRC) in industrial assembly operations is revolutionizing the manufacturing landscape. In this evolving environment, operators are required to seamlessly coordinate their manual tasks with real-time task information and robotic behaviors. These demands fluctuate during operation, yet conventional workload assessments depend on body-worn physiological sensors that complicate practical deployment. Here, we present a vision-based attention--action framework for continuous and interpretable workload-related assessment in HRC assembly. The framework combines RGB-D observations with robot states and calibrated task-related areas to construct a temporally confirmed representation of operator behavior. This representation identifies where task demand is concentrated and explains how it develops when attention and action diverge, the task context changes, or the operator hesitates. We evaluated the framework in a three-level collaborative gearbox assembly experiment with ten participants, using subjective ratings and synchronized physiological signals as independent references. Raw NASA-TLX ratings confirmed increasing perceived workload across conditions, with significant effects on overall workload and its mental and temporal dimensions. The vision-derived HRC-CWL output was significantly associated with ECG-derived features in seven of nine participants with complete correlation data. Synchronized interaction episodes further showed temporal correspondence between detected hesitation and physiological activity. Real-time deployment demonstrated that the framework can operate without requiring operators to wear additional sensors. These findings support HRC-CWL as an interpretable behavioral proxy for cognitive ergonomics analysis and adaptive robot assistance, rather than a direct psychophysiological measure of workload.

📄 PDF Abstract BibTeX arXiv:2609.15232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gaze2Segment: A Pilot Study for Integrating Eye-Tracking Technology into Medical Image Segmentation

2016-08-10 · Naji Khosravan, Haydar Celik, Baris Turkbey, Ruida Cheng 외

This study introduced a novel system, called Gaze2Segment, integrating biological and computer vision techniques to support radiologists' reading experience with an automatic image segmentation task. During diagnostic as…

DiagnosticImage SegmentationMedical Image SegmentationSegmentation+1

Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

2026-03-16 · Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui 외 arxiv

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observa…

Cooperative Dual Attention for Audio-Visual Speech Enhancement with Facial Cues

2023-11-24 · Feixiang Wang, Shuang Yang, Shiguang Shan, Xilin Chen

In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects additional speech-related attributes such…

Speech Enhancement

Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos

2025-03-21 · Yuang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song 외

Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due…

object-detectionObject Detection

MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval

2026-03-18 · Xuri Ge, Chunhao Wang, Xindi Wang, Zheyun Qin 외 arxiv

Composed Image Retrieval (CIR) aims to retrieve target images based on a reference image and modified texts. However, existing methods often struggle to extract the correct semantic cues from the reference image that bes…

Image Retrieval