paper-with-me

홈 › Papers

E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

2025-12-04 · Yihong Tang, Haicheng Liao, Tong Nie, Junlin He, Ao Qu, Kehua Chen, Wei Ma, Zhenning Li, Lijun Sun, Chengzhong Xu arxiv

End-to-end autonomous driving (AD) systems increasingly adopt vision-language-action (VLA) models, yet they typically ignore the passenger's emotional state, which is central to comfort and AD acceptance. We introduce Open-Domain End-to-End (OD-E2E) autonomous driving, where an autonomous vehicle (AV) must interpret free-form natural-language commands, infer the emotion, and plan a physically feasible trajectory. We propose E3AD, an emotion-aware VLA framework that augments semantic understanding with two cognitively inspired components: a continuous Valenc-Arousal-Dominance (VAD) emotion model that captures tone and urgency from language, and a dual-pathway spatial reasoning module that fuses egocentric and allocentric views for human-like spatial cognition. A consistency-oriented training scheme, combining modality pretraining with preference-based alignment, further enforces coherence between emotional intent and driving actions. Across real-world datasets, E3AD improves visual grounding and waypoint planning and achieves state-of-the-art (SOTA) VAD correlation for emotion estimation. These evaluation results show that injecting emotion into VLA-style driving yields more human-aligned grounding, planning, and feedback.

📄 PDF Abstract BibTeX arXiv:2512.04733

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingSpatial ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Socratis: Are large multimodal models emotionally aware?

2023-08-31 · Katherine Deng, Arijit Ray, Reuben Tan, Saadia Gabriel 외

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reasons. Learning diverse reactions to multi…

Articles

A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction

2026-01-08 · Qing Wang, Zehan Li, Yaodong Song, Hongjie Chen 외 arxiv

This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT). IEAT incorporates user emotional state…

Empathetic Response GenerationEmotional IntelligenceTrajectory Modeling

AIVA: An AI-based Virtual Companion for Emotion-aware Interaction

2025-09-03 · Chenxi Li arxiv

Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing…

Natural Language UnderstandingContrastive LearningPrompt Engineering

Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation

2025-05-14 · Feifan Wang, Tengfei Song, Minggui He, Chang Su 외

Facial emotion perception in the vision large language model (VLLM) is crucial for achieving natural human-machine interaction. However, creating high-quality annotations for both coarse- and fine-grained facial emotion …

Emotion RecognitionLarge Language Model

Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction

2026-03-09 · Xiaofeng Yu, Jiaheng Dong, Jean Honorio, Abhirup Ghosh 외 arxiv

Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguous nature of human emotional expression.…

Speech Emotion Recognition