paper-with-me

홈 › Papers

Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games

2025-10-30 · Jingran Zhang, Ning Li, Justin Cui arxiv

OpenAI's ChatGPT Atlas introduces new capabilities for web interaction, enabling the model to analyze webpages, process user intents, and execute cursor and keyboard inputs directly within the browser. While its capacity for information retrieval tasks has been demonstrated, its performance in dynamic, interactive environments remains less explored. In this study, we conduct an early evaluation of Atlas's web interaction capabilities using browser-based games as test scenarios, including Google's T-Rex Runner, Sudoku, Flappy Bird, and Stein.world. We employ in-game performance scores as quantitative metrics to assess performance across different task types. Our results show that Atlas performs strongly in logical reasoning tasks like Sudoku, completing puzzles significantly faster than human baselines, but struggles substantially in real-time games requiring precise timing and motor control, often failing to progress beyond initial obstacles. These findings suggest that while Atlas demonstrates capable analytical processing, there remain notable limitations in dynamic web environments requiring real-time interaction. The website of our project can be found at https://atlas-game-eval.github.io.

📄 PDF Abstract BibTeX arXiv:2510.26298

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalLogical Reasoning

Similar Papers 제목 키워드 기반

Emerging Frontiers: Exploring the Impact of Generative AI Platforms on University Quantitative Finance Examinations

2023-08-15 · Rama K. Malladi

This study evaluated three Artificial Intelligence (AI) large language model (LLM) enabled platforms - ChatGPT, BARD, and Bing AI - to answer an undergraduate finance exam with 20 quantitative questions across various di…

Language ModelingLanguage ModellingLarge Language Model

Exploring New Frontiers in Agricultural NLP: Investigating the Potential of Large Language Models for Food Applications

2023-06-20 · Saed Rezayi, Zhengliang Liu, Zihao Wu, Chandra Dhakal 외

This paper explores new frontiers in agricultural natural language processing by investigating the effectiveness of using food-related text corpora for pretraining transformer-based language models. In particular, we foc…

Language ModellingNutrition

Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review

2024-01-03 · Luoma Ke, Song Tong, Peng Cheng, Kaiping Peng

This paper explores the frontiers of large language models (LLMs) in psychology applications. Psychology has undergone several theoretical changes, and the current use of Artificial Intelligence (AI) and Machine Learning…

Experimental DesignText Generation

InterAct: Exploring the Potentials of ChatGPT as a Cooperative Agent

2023-08-03 · Po-Lin Chen, Cheng-Shang Chang

This research paper delves into the integration of OpenAI's ChatGPT into embodied agent systems, evaluating its influence on interactive decision-making benchmark. Drawing a parallel to the concept of people assuming rol…

Decision MakingLanguage ModelingLanguage ModellingPrompt Engineering+1

Silver-Tongued and Sundry: Exploring Intersectional Pronouns with ChatGPT

2024-05-13 · Takao Fujii, Katie Seaborn, Madeleine Steeds

ChatGPT is a conversational agent built on a large language model. Trained on a significant portion of human output, ChatGPT can mimic people to a degree. As such, we need to consider what social identities ChatGPT simul…

Language ModelingLanguage ModellingLarge Language Model