paper-with-me

Papers

TVWorld: Foundations for Remote-Control TV Agents

2026-01-19 · Zhantao Ma, Quanfeng Lu, Shuai Zhong, Dahai Yu, Ping Luo, Michael K. Ng arxiv

Recent large vision-language models (LVLMs) have demonstrated strong potential for device control. However, existing research has primarily focused on point-and-click (PnC) interaction, while remote-control (RC) interaction commonly encountered in everyday TV usage remains largely underexplored. To fill this gap, we introduce \textbf{TVWorld}, an offline graph-based abstraction of real-world TV navigation that enables reproducible and deployment-free evaluation. On this basis, we derive two complementary benchmarks that comprehensively assess TV-use capabilities: \textbf{TVWorld-N} for topology-aware navigation and \textbf{TVWorld-G} for focus-aware grounding. These benchmarks expose a key limitation of existing agents: insufficient topology awareness for focus-based, long-horizon TV navigation. Motivated by this finding, we propose a \emph{Topology-Aware Training} framework that injects topology awareness into LVLMs. Using this framework, we develop \textbf{TVTheseus}, a foundation model specialized for TV navigation. TVTheseus achieves a success rate of $68.3\%$ on TVWorld-N, surpassing strong closed-source baselines such as Gemini 3 Flash and establishing state-of-the-art (SOTA) performance. Additional analyses further provide valuable insights into the development of effective TV-use agents.

📄 PDF Abstract BibTeX arXiv:2601.13142

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Formal Foundations of Agentic Business Process Management

2026-04-19 · Giuseppe De Giacomo, Timotheus Kampik, Lukas Kirchdorfer, Marco Montali 외 arxiv

Just like traditional BPM systems, agentic BPM systems are built around a specification of the process under consideration. Their distinguishing feature, however, is that the execution of the process is driven by multipl…

Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems

2026-01-05 · Niloufar Alipour Talemi, Julia Boone, Fatemeh Afghah arxiv

The paradigm of Earth Observation analysis is shifting from static deep learning models to autonomous agentic AI. Although recent vision foundation models and multimodal large language models advance representation learn…

Representation Learning

Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework

2026-01-07 · Ben Carvell, Marc Thomas, Andrew Pace, Christopher Dorney 외 arxiv

We present a rigorous, human-in-the-loop evaluation framework for assessing the performance of AI agents on the task of Air Traffic Control, grounded in a regulator-certified simulator-based curriculum used for training …

VLAgents: A Policy Server for Efficient VLA Inference

2026-01-16 · Tobias Jülg, Khaled Gamal, Nisarga Nilavadi, Pierre Krack 외 arxiv

The rapid emergence of Vision-Language-Action models (VLAs) has a significant impact on robotics. However, their deployment remains complex due to the fragmented interfaces and the inherent communication latency in distr…

A Multi-agent Market Model Can Explain the Impact of AI Traders in Financial Markets -- A New Microfoundations of GARCH model

2024-09-19 · Kei Nakagawa, Masanori Hirano, Kentaro Minami, Takanobu Mizuta

The AI traders in financial markets have sparked significant interest in their effects on price formation mechanisms and market volatility, raising important questions for market stability and regulation. Despite this in…