paper-with-me

Papers

MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis

2024-09-11 · Hanyu Jiang, Jian Xue, Xing Lan, Guohong Hu, Ke Lu

This paper introduces MVLLaVA, an intelligent agent designed for novel view synthesis tasks. MVLLaVA integrates multiple multi-view diffusion models with a large multimodal model, LLaVA, enabling it to handle a wide range of tasks efficiently. MVLLaVA represents a versatile and unified platform that adapts to diverse input types, including a single image, a descriptive caption, or a specific change in viewing azimuth, guided by language instructions for viewpoint generation. We carefully craft task-specific instruction templates, which are subsequently used to fine-tune LLaVA. As a result, MVLLaVA acquires the capability to generate novel view images based on user instructions, demonstrating its flexibility across diverse tasks. Experiments are conducted to validate the effectiveness of MVLLaVA, demonstrating its robust performance and versatility in tackling diverse novel view synthesis challenges.

📄 PDF Abstract BibTeX arXiv:2409.07129

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveNovel View Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Gentopia: A Collaborative Platform for Tool-Augmented LLMs

2023-08-08 · Binfeng Xu, Xukun Liu, Hua Shen, Zeyu Han 외

Augmented Language Models (ALMs) empower large language models with the ability to use tools, transforming them into intelligent agents for real-world interactions. However, most existing frameworks for ALMs, to varying …

PyTSC: A Unified Platform for Multi-Agent Reinforcement Learning in Traffic Signal Control

2024-10-23 · Rohit Bokade, Xiaoning Jin

Multi-Agent Reinforcement Learning (MARL) presents a promising approach for addressing the complexity of Traffic Signal Control (TSC) in urban environments. However, existing platforms for MARL-based TSC research face ch…

ManagementMulti-agent Reinforcement LearningTraffic Signal Control

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

2026-04-20 · Haoyu Wu, Jiwen Yu, Yingtian Zou, Xihui Liu arxiv

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical fram…

Robot ManipulationVideo Generation

KG-MAS: Knowledge Graph-Enhanced Multi-Agent Infrastructure for coupling physical and digital robotic environments

2025-10-11 · Walid Abdela arxiv

The seamless integration of physical and digital environments in Cyber-Physical Systems(CPS), particularly within Industry 4.0, presents significant challenges stemming from system heterogeneity and complexity. Tradition…

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

2025-11-18 · Mingyue Cheng, Shuo Yu, Daoyu Wang, Qingchuan Li 외 arxiv

Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agents take on more complex reasoning, decision making, tool use, and long-h…

Reinforcement LearningDecision Making