paper-with-me

홈 › Papers

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

2026-03-07 · Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang arxiv

We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet over centuries with sub-meter, sub-second precision. Multi-modal encoders (e.g. vision-language models) are fused with Earth4D embeddings and trained via masked reconstruction. We demonstrate Earth4D's expressive power by achieving state-of-the-art performance on an ecological forecasting benchmark. Earth4D with learnable hash probing surpasses a multi-modal foundation model pre-trained on substantially more data. Access open source code and download models at: https://github.com/legel/deepearth

📄 PDF Abstract BibTeX arXiv:2603.07039

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Supervised Enhancement of Forward-Looking Sonar Images: Bridging Cross-Modal Degradation Gaps through Feature Space Transformation and Multi-Frame Fusion

2025-04-15 · Zhisheng Zhang, Peng Zhang, Fengxiang Wang, Liangli Ma 외

Enhancing forward-looking sonar images is critical for accurate underwater target detection. Current deep learning methods mainly rely on supervised training with simulated data, but the difficulty in obtaining high-qual…

Multi-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition

2024-04-16 · Marah Halawa, Florian Blume, Pia Bideau, Martin Maier 외

Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when desi…

Emotion ClassificationEmotion Recognition in ConversationFacial Expression RecognitionSelf-Supervised Learning

BEVWorld: A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space

2024-07-08 · Yumeng Zhang, Shi Gong, Kaixin Xiong, Xiaoqing Ye 외

World models are receiving increasing attention in autonomous driving for their ability to predict potential future scenarios. In this paper, we present BEVWorld, a novel approach that tokenizes multimodal sensor inputs …

Autonomous DrivingDecodermotion prediction

Multi-Modal Mutual Information (MuMMI) Training for Robust Self-Supervised Deep Reinforcement Learning

2021-07-06 · Kaiqi Chen, Yong Lee, Harold Soh

This work focuses on learning useful and robust deep world models using multiple, possibly unreliable, sensors. We find that current methods do not sufficiently encourage a shared representation between modalities; this …

Deep Reinforcement LearningMuJoCoreinforcement-learningReinforcement Learning (RL)

Self-Supervised Multimodal NeRF for Autonomous Driving

2025-06-24 · Gaurav Sharma, Ravi Kothari, Josef Schmid

In this paper, we propose a Neural Radiance Fields (NeRF) based framework, referred to as Novel View Synthesis Framework (NVSF). It jointly learns the implicit neural representation of space and time-varying scene for bo…

Autonomous DrivingNeRFNovel View Synthesis