paper-with-me

홈 › Papers

Multi-modal user interface control detection using cross-attention

2026-04-08 · Milad Moradi, Ke Yan, David Colwell, Matthias Samwald, Rhona Asgari arxiv

Detecting user interface (UI) controls from software screenshots is a critical task for automated testing, accessibility, and software analytics, yet it remains challenging due to visual ambiguities, design variability, and the lack of contextual cues in pixel-only approaches. In this paper, we introduce a novel multi-modal extension of YOLOv5 that integrates GPT-generated textual descriptions of UI images into the detection pipeline through cross-attention modules. By aligning visual features with semantic information derived from text embeddings, our model enables more robust and context-aware UI control detection. We evaluate the proposed framework on a large dataset of over 16,000 annotated UI screenshots spanning 23 control classes. Extensive experiments compare three fusion strategies, i.e. element-wise addition, weighted sum, and convolutional fusion, demonstrating consistent improvements over the baseline YOLOv5 model. Among these, convolutional fusion achieved the strongest performance, with significant gains in detecting semantically complex or visually ambiguous classes. These results establish that combining visual and textual modalities can substantially enhance UI element detection, particularly in edge cases where visual information alone is insufficient. Our findings open promising opportunities for more reliable and intelligent tools in software testing, accessibility support, and UI analytics, setting the stage for future research on efficient, robust, and generalizable multi-modal detection systems.

📄 PDF Abstract BibTeX arXiv:2604.06934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MIRIAM: A Multimodal Chat-Based Interface for Autonomous Systems

2018-03-06 · Helen Hastie, Francisco J. Chiyah Garcia, David A. Robb, Pedro Patron 외

We present MIRIAM (Multimodal Intelligent inteRactIon for Autonomous systeMs), a multimodal interface to support situation awareness of autonomous vehicles through chat-based interaction. The user is able to chat about t…

Autonomous Vehicles

An Emotion-based Korean Multimodal Empathetic Dialogue System

2022-10-01 · CAI (COLING) 2022 10 · Minyoung Jung, Yeongbeom Lim, San Kim, Jin Yea Jang 외

We propose a Korean multimodal dialogue system targeting emotion-based empathetic dialogues because most research in this field has been conducted in a few languages such as English and Japanese and in certain circumstan…

Designing The Drive: Enhancing User Experience through Adaptive Interfaces in Autonomous Vehicles

2025-12-14 · Reeteesha Roy arxiv

With the recent development and integration of autonomous vehicles (AVs) in transportation systems of the modern world, the emphasis on customizing user interfaces to optimize the overall user experience has been growing…

Autonomous Vehicles

VUT: Versatile UI Transformer for Multimodal Multi-Task User Interface Modeling

2021-09-29 · Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani 외

User interface modeling is inherently multimodal, which involves several distinct types of data: images, structures and language. The tasks are also diverse, including object detection, language generation and grounding.…

object-detectionObject DetectionQuestion AnsweringText Generation

VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling

2021-12-10 · Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani 외

User interface modeling is inherently multimodal, which involves several distinct types of data: images, structures and language. The tasks are also diverse, including object detection, language generation and grounding.…

object-detectionObject DetectionQuestion AnsweringText Generation