paper-with-me

Papers

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

2026-01-20 · Yuhua Jin, Nikita Kuzmin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou arxiv

Drones operating in human-occupied spaces suffer from insufficient communication mechanisms that create uncertainty about their intentions. We present HoverAI, an embodied aerial agent that integrates drone mobility, infrastructure-independent visual projection, and real-time conversational AI into a unified platform. Equipped with a MEMS laser projector, onboard semi-rigid screen, and RGB camera, HoverAI perceives users through vision and voice, responding via lip-synced avatars that adapt appearance to user demographics. The system employs a multimodal pipeline combining VAD, ASR (Whisper), LLM-based intent classification, RAG for dialogue, face analysis for personalization, and voice synthesis (XTTS v2). Evaluation demonstrates high accuracy in command recognition (F1: 0.90), demographic estimation (gender F1: 0.89, age MAE: 5.14 years), and speech transcription (WER: 0.181). By uniting aerial robotics with adaptive conversational AI and self-contained visual output, HoverAI introduces a new class of spatially-aware, socially responsive embodied agents for applications in guidance, assistance, and human-centered interaction.

📄 PDF Abstract BibTeX arXiv:2601.13801

Code (0)

등록된 구현이 없습니다.

Tasks

Intent Classification

Similar Papers 제목 키워드 기반

Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories

2025-09-19 · Yifan Lin, Sophie Ziyu Liu, Ran Qi, George Z. Xue 외 arxiv

We present Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories (ACDC), an autonomous drone cinematography system driven by natural language communication between human directors and drones. The ma…

CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

2025-05-08 · Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng 외

Aerial vision-and-language navigation (VLN), requiring drones to interpret natural language instructions and navigate complex urban environments, emerges as a critical embodied AI challenge that bridges human-robot inter…

Large Language ModelNavigateSpatial ReasoningVision and Language Navigation

AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models

2024-08-28 · Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun 외

Aerospace embodied intelligence aims to empower unmanned aerial vehicles (UAVs) and other aerospace platforms to achieve autonomous perception, cognition, and action, as well as egocentric active interaction with humans …

Spatial ReasoningTask Planning

AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding

2025-04-13 · Fei Lin, Yonglin Tian, Tengchao Zhang, Jun Huang 외

Unmanned Aerial Vehicles (UAVs) are increasingly important in dynamic environments such as logistics transportation and disaster response. However, current tasks often rely on human operators to monitor aerial videos and…

Disaster ResponseScheduling

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

2026-07-11 · Weichen Zhang, Shiquan Yu, Yinan Zhu, Peizhi Tang 외 arxiv

We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aeria…

Question Answering