paper-with-me

Papers

Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark

2024-11-29 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman, Viorica Pătrăucean

Following the successful 2023 edition, we organised the Second Perception Test challenge as a half-day workshop alongside the IEEE/CVF European Conference on Computer Vision (ECCV) 2024, with the goal of benchmarking state-of-the-art video models and measuring the progress since last year using the Perception Test benchmark. This year, the challenge had seven tracks (up from six last year) and covered low-level and high-level tasks, with language and non-language interfaces, across video, audio, and text modalities; the additional track covered hour-long video understanding and introduced a novel video QA benchmark 1h-walk VQA. Overall, the tasks in the different tracks were: object tracking, point tracking, temporal action localisation, temporal sound localisation, multiple-choice video question-answering, grounded video question-answering, and hour-long video question-answering. We summarise in this report the challenge tasks and results, and introduce in detail the novel hour-long video QA benchmark 1h-walk VQA.

📄 PDF Abstract BibTeX arXiv:2411.19941

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject TrackingPoint TrackingQuestion AnsweringVideo Question AnsweringVideo UnderstandingVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Perception Test 2025: Challenge Summary and a Unified VQA Extension

2026-01-09 · Joseph Heyward, Nikhil Parthasarathy, Tyler Zhu, Aravindh Mahendran 외 arxiv

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and …

Video GenerationObject TrackingPoint Tracking

Perception Test 2023: A Summary of the First Challenge And Outcome

2023-12-20 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking state-of-the-art video models on the recen…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+3

Long-Term Identity-Aware Multi-Person Tracking for Surveillance Video Summarization

2016-04-25 · Shoou-I Yu, Yi Yang, Xuanchong Li, Alexander G. Hauptmann

Multi-person tracking plays a critical role in the analysis of surveillance video. However, most existing work focus on shorter-term (e.g. minute-long or hour-long) video sequences. Therefore, we propose a multi-person t…

Face RecognitionVideo Summarization

Vision-Centric BEV Perception: A Survey

2022-08-04 · Yuexin Ma, Tai Wang, Xuyang Bai, Huitong Yang 외

In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the worl…

Survey

HourVideo: 1-Hour Video-Language Understanding

2024-11-07 · Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota 외

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, tempora…

BenchmarkingcounterfactualMultiple-choiceRetrieval+1