paper-with-me

Papers

Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval

2024-06-22 · Paul Primus, Gerhard Widmer

Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that utilizes audio metadata as an additional clue to understand the content of audio signals before matching them with textual queries. We experimented with metadata often attached to audio recordings, such as keywords and natural-language descriptions, and we investigated late and mid-level fusion strategies to merge audio and metadata. Our hybrid approach with keyword metadata and late fusion improved the retrieval performance over a content-based baseline by 2.36 and 3.69 pp. mAP@10 on the ClothoV2 and AudioCaps benchmarks, respectively.

📄 PDF Abstract BibTeX arXiv:2406.15897

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsRetrieval

Similar Papers 제목 키워드 기반

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

2025-12-08 · Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting, Patrick Schramowski 외 arxiv

Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is commonly used to create and curate these dat…

AutoML using Metadata Language Embeddings

2019-10-08 · Iddo Drori, Lu Liu, Yi Nian, Sharath C. Koorathota 외

As a human choosing a supervised learning algorithm, it is natural to begin by reading a text description of the dataset and documentation for the algorithms you might use. We demonstrate that the same idea improves the …

AutoML

Listen, Read, and Identify: Multimodal Singing Language Identification of Music

2021-03-02 · Keunwoo Choi, Yuxuan Wang

We propose a multimodal singing language classification model that uses both audio content and textual metadata. LRID-Net, the proposed model, takes an audio signal and a language probability vector estimated from the me…

Language Identification

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

2026-05-30 · Shenghao Ding arxiv

This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around …

LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting

2025-05-29 · Pai Zhu, Quan Wang, Dhruuv Agarwal, Kurt Partridge

Custom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models ar…

Keyword Spottingtext-to-speechText to Speech