paper-with-me

홈 › Papers

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

2026-08-30 · Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong hf

GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they are repeatedly reused for future interaction. We introduce GUI-CC, a benchmark that evaluates contextual consistency of GUI world models as agent environments rather than isolated next-screen predictors. GUI-CC contains two complementary tracks: an offline reference-action track that rolls models along real mobile GUI trajectories, and an online agent-loop track that lets fixed probing agents interact with model-generated UIs. We construct 500 offline trajectory tasks from GUIOdyssey and 200 emulator-verified online tasks across 30 mobile apps. GUI-CC evaluates transition fidelity, transition plausibility, contextual consistency, and task progress. Experiments show that plausible single-step generation does not guarantee reliable environment simulation: current models often produce usable-looking screens while failing to preserve task-relevant context or support executable multi-step rollouts.

📄 PDF Abstract BibTeX arXiv:2609.00048

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MIND: Benchmarking Memory Consistency and Action Control in World Models

2026-02-08 · Yixuan Ye, Xuanyu Lu, Yuxin Jiang, Yuchao Gu 외 arxiv

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first ope…

Simulated Contextual Bandits for Personalization Tasks from Recommendation Datasets

2022-10-12 · Anton Dereventsov, Anton Bibin

We propose a method for generating simulated contextual bandit environments for personalization tasks from recommendation datasets like MovieLens, Netflix, Last.fm, Million Song, etc. This allows for personalization envi…

BenchmarkingMulti-Armed Bandits

Benchmarking Contextual Understanding for In-Car Conversational Systems

2025-12-12 · Philipp Habicht, Lev Sorokin, Abdullah Saydemir, Ken E. Friedl 외 arxiv

In-Car Conversational Question Answering (ConvQA) systems significantly enhance user experience by enabling seamless voice interactions. However, assessing their accuracy and reliability remains a challenge. This paper e…

Conversational Question Answering

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

2026-01-12 · Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang 외 arxiv

With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user trust and engagement. However, existing …

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

2026-05-25 · Yusong Lin, Xinyuan Liang, Haiyang Wang, Qipeng Gu 외 arxiv

Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, li…