DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types.
Code (1)
Tasks
Code GenerationSynthetic Data GenerationSimilar Papers 제목 키워드 기반
Telling Stories from Computational Notebooks: AI-Assisted Presentation Slides Creation for Presenting Data Science Work
Creating presentation slides is a critical but time-consuming task for data scientists. While researchers have proposed many AI techniques to lift data scientists' burden on data preparation and model selection, few have…
Model SelectionVisual grounding for desktop graphical user interfaces
Most instance perception and image understanding solutions focus mainly on natural images. However, applications for synthetic images, and more specifically, images of Graphical User Interfaces (GUI) remain limited. This…
Language ModelingLanguage ModellingLarge Language Modelobject-detection+2Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant
Explainable artificial intelligence (XAI) methods are being proposed to help interpret and understand how AI systems reach specific predictions. Inspired by prior work on conversational user interfaces, we argue that aug…
AllDecision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+3FedWSIDD: Federated Whole Slide Image Classification via Dataset Distillation
Federated learning (FL) has emerged as a promising approach for collaborative medical image analysis, enabling multiple institutions to build robust predictive models while preserving sensitive patient data. In the conte…
ClassificationDataset DistillationFederated Learningimage-classification+2RecGaze: The First Eye Tracking and User Interaction Dataset for Carousel Interfaces
Carousel interfaces are widely used in e-commerce and streaming services, but little research has been devoted to them. Previous studies of interfaces for presenting search and recommendation results have focused on sing…
Recommendation Systems