paper-with-me

홈 › Papers

EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

2024-06-14 · Julian Straub, Daniel DeTone, Tianwei Shen, Nan Yang, Chris Sweeney, Richard Newcombe

The advent of wearable computers enables a new source of context for AI that is embedded in egocentric sensor data. This new egocentric data comes equipped with fine-grained 3D location information and thus presents the opportunity for a novel class of spatial foundation models that are rooted in 3D space. To measure progress on what we term Egocentric Foundation Models (EFMs) we establish EFM3D, a benchmark with two core 3D egocentric perception tasks. EFM3D is the first benchmark for 3D object detection and surface regression on high quality annotated egocentric data of Project Aria. We propose Egocentric Voxel Lifting (EVL), a baseline for 3D EFMs. EVL leverages all available egocentric modalities and inherits foundational capabilities from 2D foundation models. This model, trained on a large simulated dataset, outperforms existing methods on the EFM3D benchmark.

📄 PDF Abstract BibTeX arXiv:2406.10224

Code (1)

facebookresearch/efm3d 공식 구현 pytorch

Tasks

3D Object Detection3D ReconstructionMulti-View 3D Reconstructionobject-detectionObject Detection

Similar Papers 제목 키워드 기반

EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

2025-05-16 · Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu 외

Imitation learning for manipulation has a well-known data scarcity problem. Unlike natural language and 2D computer vision, there is no Internet-scale corpus of data for dexterous manipulation. One appealing option is eg…

Imitation LearningTrajectory Prediction

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

2026-05-18 · Dongyan Lin, Phillip Rust, Angel Villar Corrales, Alvin W. M. Tan 외 arxiv

Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs)…

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

2026-03-10 · Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu 외 arxiv

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from …

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

2026-07-01 · Jihyeok Jung, Jeewu Lee, Sanghyeop Kim, Chanhee Han 외 arxiv

Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-…

Scene Understanding

Does Progress On Object Recognition Benchmarks Improve Real-World Generalization?

2023-07-24 · Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark Ibrahim

For more than a decade, researchers have measured progress in object recognition on ImageNet-based generalization benchmarks such as ImageNet-A, -C, and -R. Recent advances in foundation models, trained on orders of magn…

Object Recognition