paper-with-me

홈 › Papers

Qualitative Evaluation of LLM-Designed GUI

2026-01-30 · Bartosz Sawicki, Tomasz Les, Dariusz Parzych, Aleksandra Wycisk-Ficek, Pawel Trebacz, Pawel Zawadzki arxiv

As generative artificial intelligence advances, Large Language Models (LLMs) are being explored for automated graphical user interface (GUI) design. This study investigates the usability and adaptability of LLM-generated interfaces by analysing their ability to meet diverse user needs. The experiments included utilization of three state-of-the-art models from January 2025 (OpenAI GPT o3-mini-high, DeepSeek R1, and Anthropic Claude 3.5 Sonnet) generating mockups for three interface types: a chat system, a technical team panel, and a manager dashboard. Expert evaluations revealed that while LLMs are effective at creating structured layouts, they face challenges in meeting accessibility standards and providing interactive functionality. Further testing showed that LLMs could partially tailor interfaces for different user personas but lacked deeper contextual understanding. The results suggest that while LLMs are promising tools for early-stage UI prototyping, human intervention remains critical to ensure usability, accessibility, and user satisfaction.

📄 PDF Abstract BibTeX arXiv:2601.22759

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

2026-05-17 · Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare, Mohammad Sabouri arxiv

The rapid progress of generative artificial intelligence has exposed fundamental limitations in existing evaluation methodologies, particularly for open-ended, creative, and human-facing tasks. Traditional automatic metr…

Image Generation

Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values

2025-05-14 · Martín Rodríguez, Gustavo Rossi, Alejandro Fernandez

The design and implementation of unit tests is a complex task many programmers neglect. This research evaluates the potential of Large Language Models (LLMs) in automatically generating test cases, comparing them with ma…

AlignNet: A Unifying Approach to Audio-Visual Alignment

2020-02-12 · Jianren Wang, Zhaoyuan Fang, Hang Zhao

We present AlignNet, a model that synchronizes videos with reference audios under non-uniform and irregular misalignments. AlignNet learns the end-to-end dense correspondence between each frame of a video and an audio. O…

LITE: LLM-Impelled efficient Taxonomy Evaluation

2025-04-02 · Lin Zhang, Zhouhong Gu, SuHang Zheng, Tao Wang 외

This paper presents LITE, an LLM-based evaluation method designed for efficient and flexible assessment of taxonomy quality. To address challenges in large-scale taxonomy evaluation, such as efficiency, fairness, and con…

Fairness

Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale

2026-03-09 · Daryl Hedley, Doug Pietrzak, Jorge Dias, Ian Burden 외 arxiv

Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and instructional processes. However, traditio…