paper-with-me

Papers

Replay: Multi-modal Multi-view Acted Videos for Casual Holography

2023-07-22 · ICCV 2023 1 · Roman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotny, Andrea Vedaldi, Changan Chen, Filippos Kokkinos, Ben Graham, Natalia Neverova

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras, and recorded with a large array of microphones at different positions in the room. Overall, the dataset contains over 4000 minutes of footage and over 7 million timestamped high-resolution frames annotated with camera poses and partially with foreground masks. The Replay dataset has many potential applications, such as novel-view synthesis, 3D reconstruction, novel-view acoustic synthesis, human body and face analysis, and training generative models. We provide a benchmark for training and evaluating novel-view synthesis, with two scenarios of different difficulty. Finally, we evaluate several baseline state-of-the-art methods on the new benchmark.

📄 PDF Abstract BibTeX arXiv:2307.12067

Code (1)

facebookresearch/replay_dataset 공식 구현 pytorch

Tasks

3D ReconstructionNovel View Synthesis

Similar Papers 제목 키워드 기반

Towards Universal Soccer Video Understanding

2024-12-02 · CVPR 2025 1 · Jiayuan Rao, HaoNing Wu, Hao Jiang, Ya zhang 외

As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we mak…

Action ClassificationSports UnderstandingVideo Understanding

Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks

2025-01-09 · Seyed Amir Bidaki, Amir Mohammadkhah, Kiyan Rezaee, Faeze Hassani 외

Online Continual Learning (OCL) is a critical area in machine learning, focusing on enabling models to adapt to evolving data streams in real-time while addressing challenges such as catastrophic forgetting and the stabi…

Continual Learningimage-classificationImage Classificationobject-detection+3

A Review Paper of the Effects of Distinct Modalities and ML Techniques to Distracted Driving Detection

2025-01-20 · Anthony. Dontoh, Stephanie. Ivey, Logan. Sirbaugh, Armstrong. Aboah

Distracted driving remains a significant global challenge with severe human and economic repercussions, demanding improved detection and intervention strategies. While previous studies have extensively explored single-mo…

Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning

2025-11-17 · Xinlan Wu, Bin Zhu, Feng Han, Pengkun Jiao 외 arxiv

Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LMMs) in food analysis suffer from catastr…

Semantic SimilarityContinual Learning

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

2026-06-22 · Chuangxin Zhao, Canran Xiao, Siyuan Ma, Mengyao Lyu 외 arxiv

Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-tuning often causes severe forgetting of …

Continual Learning