paper-with-me

홈 › Papers

VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft

2025-08-26 · Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, Hao Wang arxiv

Large language models (LLMs) have shown significant promise in embodied decision-making tasks within virtual open-world environments. Nonetheless, their performance is hindered by the absence of domain-specific knowledge. Methods that finetune on large-scale domain-specific data entail prohibitive development costs. This paper introduces VistaWise, a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis. It reduces the requirement for domain-specific training data from millions of samples to a few hundred. VistaWise integrates visual information and textual dependencies into a cross-modal knowledge graph (KG), enabling a comprehensive and accurate understanding of multimodal environments. We also equip the agent with a retrieval-based pooling strategy to extract task-related information from the KG, and a desktop-level skill library to support direct operation of the Minecraft desktop client via mouse and keyboard inputs. Experimental results demonstrate that VistaWise achieves state-of-the-art performance across various open-world tasks, highlighting its effectiveness in reducing development costs while enhancing agent performance.

📄 PDF Abstract BibTeX arXiv:2508.18722

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

STEMS: Spatial-Temporal Enhanced Safe Multi-Agent Coordination for Building Energy Management

2025-10-15 · Huiliang Zhang, Di Wu, Arnaud Zinflou, Benoit Boulet arxiv

Building energy management is essential for achieving carbon reduction goals, improving occupant comfort, and reducing energy costs. Coordinated building energy management faces critical challenges in exploiting spatial-…

Multi-agent Reinforcement LearningGraph Representation Learning

Adaptive In-conversation Team Building for Language Model Agents

2024-05-29 · Linxin Song, Jiale Liu, Jieyu Zhang, Shaokun Zhang 외

Leveraging multiple large language model (LLM) agents has shown to be a promising approach for tackling complex tasks, while the effective design of multiple agents for a particular application remains an art. It is thus…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

AgentBalance: Backbone-then-Topology Design for Cost-Effective Multi-Agent Systems under Budget Constraints

2025-12-12 · Shuowei Cai, Yansong Ning, Hao Liu arxiv

Large Language Model (LLM)-based multi-agent systems (MAS) are becoming indispensable building blocks for web-scale applications such as web search, social network analytics, and online customer support, where cost-effec…

Representation Learning

Efficient Agents: Building Effective Agents While Reducing Cost

2025-07-24 · Ningning Wang, Xavier Hu, Pai Liu, He Zhu 외 arxiv

The remarkable capabilities of Large Language Model (LLM)-driven agents have enabled sophisticated systems to tackle complex, multi-step tasks, but their escalating costs threaten scalability and accessibility. This work…

TongUI: Building Generalized GUI Agents by Learning from Multimodal Web Tutorials

2025-04-17 · Bofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang 외

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing g…

Articles