paper-with-me

Papers

Large Scale Audio-Visual Video Analytics Platform for Forensic Investigations of Terroristic Attacks

2018-11-28 · Alexander Schindler, Martin Boyer, Andrew Lindley, David Schreiber, Thomas Philipp

The forensic investigation of a terrorist attack poses a huge challenge to the investigative authorities, as several thousand hours of video footage need to be spotted. To assist law enforcement agencies (LEA) in identifying suspects and securing evidences, we present a platform which fuses information of surveillance cameras and video uploads from eyewitnesses. The platform integrates analytical modules for different input-modalities on a scalable architecture. Videos are analyzed according their acoustic and visual content. Specifically, Audio Event Detection is applied to index the content according to attack-specific acoustic concepts. Audio similarity search is utilized to identify similar video sequences recorded from different perspectives. Visual object detection and tracking are used to index the content according to relevant concepts. The heterogeneous results of the analytical modules are fused into a distributed index of visual and acoustic concepts to facilitate rapid start of investigations, following traits and investigating witness reports.

📄 PDF Abstract BibTeX arXiv:1811.11623

Code (0)

등록된 구현이 없습니다.

Tasks

Event Detectionobject-detectionObject Detection

Similar Papers 제목 키워드 기반

On the Utility of Audiovisual Dialog Technologies and Signal Analytics for Real-time Remote Monitoring of Depression Biomarkers

2020-07-01 · WS 2020 7 · Michael Neumann, Oliver Roessler, David Suendermann-Oeft, Vikram Ramanarayanan

We investigate the utility of audiovisual dialog systems combined with speech and video analytics for real-time remote monitoring of depression at scale in uncontrolled environment settings. We collected audiovisual conv…

Aligned Better, Listen Better for Audio-Visual Large Language Models

2025-04-02 · Yuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao 외

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter…

Video Understanding

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

2023-03-22 · CVPR 2023 1 · Tiantian Geng, Teng Wang, Jinming Duan, Runmin Cong 외

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual event…

audio-visual event localization

Detection of Audio-Video Synchronization Errors Via Event Detection

2021-04-20 · Joshua P. Ebenezer, Yongjun Wu, Hai Wei, Sriram Sethuraman 외

We present a new method and a large-scale database to detect audio-video synchronization(A/V sync) errors in tennis videos. A deep network is trained to detect the visual signature of the tennis ball being hit by the rac…

Event DetectionVideo Synchronization

Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

2023-08-18 · PLOS ONE 2024 4 · Michael Joannou, Pia Rotshtein, Uta Noppeney

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Tim…