paper-with-me

Papers

Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics

2024-10-10 · Junyi Cao, Shanyan Guan, Yanhao Ge, Wei Li, Xiaokang Yang, Chao Ma

While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure neural-network-based simulators (black box), which may violate physical laws, or traditional physical simulators (white box), which rely on expert-defined equations that may not fully capture actual dynamics. We propose the Neural Material Adaptor (NeuMA), which integrates existing physical laws with learned corrections, facilitating accurate learning of actual dynamics while maintaining the generalizability and interpretability of physical priors. Additionally, we propose Particle-GS, a particle-driven 3D Gaussian Splatting variant that bridges simulation and observed images, allowing back-propagate image gradients to optimize the simulator. Comprehensive experiments on various dynamics in terms of grounded particle accuracy, dynamic rendering quality, and generalization ability demonstrate that NeuMA can accurately capture intrinsic dynamics.

📄 PDF Abstract BibTeX arXiv:2410.08257

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Visual Grounding of Learned Physical Models

2020-04-28 · ICML 2020 1 · Yunzhu Li, Toru Lin, Kexin Yi, Daniel M. Bear 외

Humans intuitively recognize objects' physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environme…

Visual Grounding

IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities

2024-08-23 · Bin Wang, Chunyu Xie, Dawei Leng, Yuhui Yin

In the field of multimodal large language models (MLLMs), common methods typically involve unfreezing the language model during training to foster profound visual understanding. However, the fine-tuning of such models wi…

Language ModelingLanguage ModellingLarge Language ModelVisual Grounding

Mamba-Adaptor: State Space Model Adaptor for Visual Recognition

2025-05-19 · CVPR 2025 1 · Fei Xie, Jiahao Nie, Yujin Tang, Wenkang Zhang 외

Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior per…

Inductive BiasMambaState Space ModelsTransfer Learning

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

2026-05-10 · Can Li, Zhoujian Li, Ren Li, Jie Gu 외 arxiv

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challeng…

LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting

2024-11-29 · CVPR 2025 1 · Xiaoyan Xing, Konrad Groh, Sezer Karaoglu, Theo Gevers 외

We introduce LumiNet, a novel architecture that leverages generative models and latent intrinsic representations for effective lighting transfer. Given a source image and a target lighting image, LumiNet synthesizes a re…