TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
Recently, image-based Large Multimodal Models (LMMs) have made significant progress in video question-answering (VideoQA) using a frame-wise approach by leveraging large-scale pretraining in a zero-shot manner. Nevertheless, these models need to be capable of finding relevant information, extracting it, and answering the question simultaneously. Currently, existing methods perform all of these steps in a single pass without being able to adapt if insufficient or incorrect information is collected. To overcome this, we introduce a modular multi-LMM agent framework based on several agents with different roles, instructed by a Planner agent that updates its instructions using shared feedback from the other agents. Specifically, we propose TraveLER, a method that can create a plan to "Traverse" through the video, ask questions about individual frames to "Locate" and store key information, and then "Evaluate" if there is enough information to answer the question. Finally, if there is not enough information, our method is able to "Replan" based on its collected knowledge. Through extensive experiments, we find that the proposed TraveLER approach improves performance on several VideoQA benchmarks without the need to fine-tune on specific datasets. Our code is available at https://github.com/traveler-framework/TraveLER.
Code (1)
Tasks
Question AnsweringVideo Question AnsweringZero-Shot Video Question AnswerMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
Effective modeling of how human travelers learn and adjust their travel behavior from interacting with transportation systems is critical for system assessment and planning. However, this task is also difficult due to th…
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make compl…
LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling
Large language models (LLMs) are increasingly used as behavioral proxies for self-interested travelers in agent-based traffic models. Although more flexible and generalizable than conventional models, the practical use o…
Reinforcement LearningAI-Driven Day-to-Day Route Choice
Understanding travelers' route choices can help policymakers devise optimal operational and planning strategies for both normal and abnormal circumstances. However, existing choice modeling methods often rely on predefin…
Decision MakingReinforcement Learning (RL)Potential destination discovery for low predictability individuals based on knowledge graph
Travelers may travel to locations they have never visited, which we call potential destinations of them. Especially under a very limited observation, travelers tend to show random movement patterns and usually have a lar…
Graph EmbeddingKnowledge Graph Embeddingvehicle detection