paper-with-me

홈 › Papers

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

2025-10-07 · Deheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao, Tianwen Qian, Xu Zheng, Guolei Sun, Ajad Chhatkuli, Xuanjing Huang, Yu-Gang Jiang, Luc Van Gool, Danda Pani Paudel arxiv

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day-night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day-night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains. The code and data can be found at https://github.com/dehezhang2/EgoNight.

📄 PDF Abstract BibTeX arXiv:2510.06218

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringDepth Estimation

Similar Papers 제목 키워드 기반

Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset

2025-06-04 · ZiRui Wang, Wenjing Bian, Xinghui Li, Yifu Tao 외

We introduce Oxford Day-and-Night, a large-scale, egocentric dataset for novel view synthesis (NVS) and visual relocalisation under challenging lighting conditions. Existing datasets often lack crucial combinations of fe…

3D geometryBenchmarkingNovel View Synthesis

Challenges and Trends in Egocentric Vision: A Survey

2025-03-19 · Xiang Li, Heqian Qiu, Lanxiao Wang, Hanwen Zhang 외

With the rapid development of artificial intelligence technologies and wearable devices, egocentric vision understanding has emerged as a new and challenging research direction, gradually attracting widespread attention …

Survey

Is Tracking really more challenging in First Person Egocentric Vision?

2025-07-21 · Matteo Dunnhofer, Zaira Manigrasso, Christian Micheloni arxiv

Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person ego…

Visual Object Tracking

Video Registration in Egocentric Vision under Day and Night Illumination Changes

2016-07-28 · Stefano Alletto, Giuseppe Serra, Rita Cucchiara

With the spread of wearable devices and head mounted cameras, a wide range of application requiring precise user localization is now possible. In this paper we propose to treat the problem of obtaining the user position …

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

2024-06-19 · Alessandro Suglia, Claudio Greco, Katie Baker, Jose L. Part 외

AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, n…

Question AnsweringSpatial ReasoningVideo CaptioningVideo Question Answering+1