paper-with-me

Papers

Multi-Source Transformer Architectures for Audiovisual Scene Classification

2022-10-18 · Wim Boes, Hugo Van hamme

In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essentially multi-source transformers employing a combination of auditory and visual features to make predictions. These models are evaluated utilizing the macro-averaged multi-class cross-entropy and accuracy metrics. In terms of the macro-averaged multi-class cross-entropy, our best model achieved a score of 0.620 on the validation data. This is slightly better than the performance of the baseline system (0.658). With regard to the accuracy measure, our best model achieved a score of 77.1\% on the validation data, which is about the same as the performance obtained by the baseline system (77.0\%).

📄 PDF Abstract BibTeX arXiv:2210.10212

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationScene Classification

Similar Papers 제목 키워드 기반

Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events

2019-12-02 · Wim Boes, Hugo Van hamme

We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a…

General ClassificationMachine TranslationTranslation

LTX-2: Efficient Joint Audio-Visual Foundation Model

2026-01-06 · Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman 외 arxiv

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source found…

Audio GenerationVideo Generation

Audiovisual Database with 360 Video and Higher-Order Ambisonics Audio for Perception, Cognition, Behavior, and QoE Evaluation Research

2022-12-27 · Thomas Robotham, Ashutosh Singla, Olli S. Rummukainen, Alexander Raake 외

Research into multi-modal perception, human cognition, behavior, and attention can benefit from high-fidelity content that may recreate real-life-like scenes when rendered on head-mounted displays. Moreover, aspects of a…

SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification

2025-06-21 · Gnana Praveen Rajasekhar, Jahangir Alam

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address th…

Contrastive LearningSelf-Supervised LearningSpeaker Verification

OAVA: the open audio-visual archives aggregator

2023-12-16 · International Journal on Digital Libraries 2023 12 · Polychronis Charitidis, Sotirios Moschos, Chrysostomos Bakouras, Stavros Doropoulos 외

The purpose of the current article is to provide an overview of an open-access audiovisual aggregation and search service platform developed for Greek audiovisual content during the OAVA (Open Access AudioVisual Archive)…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1