paper-with-me

홈 › Papers

DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos

2025-03-11 · CVPR 2025 1 · Lorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-Cantin

Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by dynamic interactions and a strong dependence on the wearer engagement with the environment. Traditional approaches often focus on isolated clips or fail to integrate rich semantic and geometric information, limiting scene comprehension. We introduce Dynamic Image-Video Feature Fields (DIV FF), a framework that decomposes the egocentric scene into persistent, dynamic, and actor based components while integrating both image and video language features. Our model enables detailed segmentation, captures affordances, understands the surroundings and maintains consistent understanding over time. DIV-FF outperforms state-of-the-art methods, particularly in dynamically evolving scenarios, demonstrating its potential to advance long term, spatio temporal scene understanding.

📄 PDF Abstract BibTeX arXiv:2503.08344

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams

2023-12-03 · CVPR 2024 1 · Liao Wang, Kaixin Yao, Chengcheng Guo, Zhirui Zhang 외

Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However, rendering dynamic, long-duration radiance fields on ubiquitous devices remains challenging, due to data storage and computation…

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

2025-03-13 · CVPR 2025 1 · Wanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song 외

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian repre…

Large Language ModelObjectSentence Embeddings

D$^2$NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

2022-05-31 · Tianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, Forrester Cole 외

Given a monocular video, segmenting and decoupling dynamic objects while recovering the static environment is a widely studied problem in machine intelligence. Existing solutions usually approach this problem in the imag…

Image SegmentationNeRFSemantic SegmentationShadow Removal

Towards Efficient Neural Scene Graphs by Learning Consistency Fields

2022-10-09 · Yeji Song, Chaerin Kong, Seoyoung Lee, Nojun Kwak 외

Neural Radiance Fields (NeRF) achieves photo-realistic image rendering from novel views, and the Neural Scene Graphs (NSG) \cite{ost2021neural} extends it to dynamic scenes (video) with multiple objects. Nevertheless, co…

NeRF

Resolution-Agnostic Neural Compression for High-Fidelity Portrait Video Conferencing via Implicit Radiance Fields

2024-02-26 · Yifei Li, Xiaohong Liu, Yicong Peng, Guangtao Zhai 외

Video conferencing has caught much more attention recently. High fidelity and low bandwidth are two major objectives of video compression for video conferencing applications. Most pioneering methods rely on classic video…

Video Compression