paper-with-me

홈 › Papers

Pretrained Conformers for Audio Fingerprinting and Retrieval

2025-08-15 · Kemal Altwlkany, Elmedin Selmanovic, Sead Delalic arxiv

Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.

📄 PDF Abstract BibTeX arXiv:2508.11609

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Robust Neural Audio Fingerprinting using Music Foundation Models

2025-11-07 · Shubhr Singh, Kiran Bhat, Xavier Riley, Benjamin Resnick 외 arxiv

The proliferation of distorted, compressed, and manipulated music on modern media platforms like TikTok motivates the development of more robust audio fingerprinting techniques to identify the sources of musical recordin…

Data Augmentation

Variable-Length Audio Fingerprinting

2026-03-25 · Hongjie Chen, Hanyu Meng, Huimin Zeng, Ryan A. Rossi 외 arxiv

Audio fingerprinting converts audio to much lower-dimensional representations, allowing distorted recordings to still be recognized as their originals through similar fingerprints. Existing deep learning approaches rigid…

Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance

2026-01-25 · Ziling Gong, Yunyan Ouyang, Iram Kamdar, Melody Ma 외 arxiv

Audio fingerprinting provides an identifiable representation of acoustic signals, which can be later used for identification and retrieval systems. To obtain a discriminative representation, the input audio is usually se…

Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization

2024-10-29 · Kemal Altwlkany, Sead Delalić, Adis Alihodžić, Elmedin Selmanović 외

Audio fingerprinting techniques have seen great advances in recent years, enabling accurate and fast audio retrieval even in conditions when the queried audio sample has been highly deteriorated or recorded in noisy cond…

GPURetrievalSpeech-to-Text

Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning

2020-10-22 · Sungkyun Chang, Donmoon Lee, Jeongsoo Park, Hyungui Lim 외

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and cou…

Audio FingerprintContrastive LearningMusic Information RetrievalRetrieval