paper-with-me

홈 › Papers

Aria-UI: Visual Grounding for GUI Instructions

2024-12-20 · Yuhao Yang, Yue Wang, Dongxu Li, Ziyang Luo, Bei Chen, Chao Huang, Junnan Li

Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce Aria-UI, a large multimodal model specifically designed for GUI grounding. Aria-UI adopts a pure-vision approach, eschewing reliance on auxiliary inputs. To adapt to heterogeneous planning instructions, we propose a scalable data pipeline that synthesizes diverse and high-quality instruction samples for grounding. To handle dynamic contexts in task performing, Aria-UI incorporates textual and text-image interleaved action histories, enabling robust context-aware reasoning for grounding. Aria-UI sets new state-of-the-art results across offline and online agent benchmarks, outperforming both vision-only and AXTree-reliant baselines. We release all training data and model checkpoints to foster further research at https://ariaui.github.io.

📄 PDF Abstract BibTeX arXiv:2412.16256

Code (1)

ariaui/aria-ui pytorch

Tasks

Natural Language Visual GroundingVisual Grounding

Similar Papers 제목 키워드 기반

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

2026-07-06 · Adrian Szvoren, Dimitrios Kanoulas, Nilufer Tuptuk arxiv

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first emp…

Semantic SegmentationRobot NavigationVisual Grounding

Layover or Direct Flight: Rethinking Audio-Guided Image Segmentation

2025-11-27 · Joel Alberto Santos, Zongwei Wu, Xavier Alameda-Pineda, Radu Timofte arxiv

Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on ve…

Image Segmentation

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

2025-03-28 · Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan 외

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-gra…

3D visual groundingAutonomous DrivingScene UnderstandingVisual Grounding

Cross-Task Knowledge Transfer for Visually-Grounded Navigation

2019-05-01 · ICLR 2019 5 · Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh 외

Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for two different tasks: learning to follow navigational instruction…

Deep Reinforcement LearningDisentanglementEmbodied Question AnsweringQuestion Answering+3

Image Difference Grounding with Natural Language

2025-04-02 · Wenxuan Wang, Zijia Zhao, Yisi Zhang, Yepeng Tang 외

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretations. This limits their applicability in…

Visual Grounding