paper-with-me

Papers

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

2025-09-30 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda, Teruko Mitamura arxiv

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite its potential use cases, the system development tailored for such an assistant is still underexplored. In this paper, we propose a novel framework, called TAMA, a Tool-Augmented Multimodal Agent, for procedural activity understanding. TAMA enables interleaved multimodal reasoning by making use of multimedia-returning tools in a training-free setting. Our experimental result on the multimodal procedural QA dataset, ProMQA-Assembly, shows that our approach can improve the performance of vision-language models, especially GPT-5 and MiMo-VL. Furthermore, our ablation studies provide empirical support for the effectiveness of two features that characterize our framework, multimedia-returning tools and agentic flexible tool selection. We believe our proposed framework and experimental results facilitate the thinking with images paradigm for video and multimodal tasks, let alone the development of procedural activity assistants.

📄 PDF Abstract BibTeX arXiv:2510.00161

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation

2025-10-06 · Dongge Han, Camille Couturier, Daniel Madrigal Diaz, Xuchao Zhang 외 arxiv

We introduce LEGOMem, a modular procedural memory framework for multi-agent large language model (LLM) systems in workflow automation. LEGOMem decomposes past task trajectories into reusable memory units and flexibly all…

YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks

2025-01-16 · Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal, Bhanu Guda 외

Multimodal AI Agents are AI models that have the capability of interactively and cooperatively assisting human users to solve day-to-day tasks. Augmented Reality (AR) head worn devices can uniquely improve the user exper…

AI AgentScene UnderstandingSSIM

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory

2025-12-08 · Sijia Li, Yuchen Huang, Zifan Liu, Zijian Li 외 arxiv

As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing approaches remain limited: full trajectories …

Reinforcement Learning

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

2024-12-16 · Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar 외

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we p…

InformativenessLarge Language ModelText GenerationText-to-Video Generation+2

TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems

2025-11-07 · Ishan Kavathekar, Hemang Jain, Ameya Rathod, Ponnurangam Kumaraguru 외 arxiv

Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents through tool use, planning, and decision-making abilities, leading to their widespread adoption across diverse tasks. As task comple…