paper-with-me

홈 › Papers

Language-Based Audio Retrieval with Converging Tied Layers and Contrastive Loss

2022-06-29 · Andrew Koh, Eng Siong Chng

In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show that using this architecture along with contrastive loss allows the model to significantly beat the performance of the baseline model. Finally, in addition to having an extremely low training memory requirement, we are able to use pretrained models as it is without needing to finetune them. We test our methods and show that using a combination of our methods beats the baseline scores significantly.

📄 PDF Abstract BibTeX arXiv:2206.14659

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Spoken Conversational Agents with Large Language Models

2025-12-02 · Chao-Han Huck Yang, Andreas Stolcke, Larry Heck arxiv

Spoken conversational agents are converging toward voice-native LLMs. This tutorial distills the path from cascaded ASR/NLU to end-to-end, retrieval-and vision-grounded systems. We frame adaptation of text LLMs to audio,…

Memory Retrieval in Transformers: Insights from The Encoding Specificity Principle

2026-01-28 · Viet Hung Dinh, Ming Ding, Youyang Qu, Kanchana Thilakarathna arxiv

While explainable artificial intelligence (XAI) for large language models (LLMs) remains an evolving field with many unresolved questions, increasing regulatory pressures have spurred interest in its role in ensuring tra…

Deep Cross-Modal Correlation Learning for Audio and Lyrics in Music Retrieval

2017-11-29 · Yu Yi, Tang Suhua, Raposo Francisco, Chen Lei

Little research focuses on cross-modal correlation learning where temporal structures of different data modalities such as audio and lyrics are taken into account. Stemming from the characteristic of temporal structures …

Retrieval

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

2026-09-10 · Haojun Zhang, Yi Zou, Min Chen, Qize Yu 외 hf

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence er…

Can Neural Network Memorization Be Localized?

2023-07-18 · Pratyush Maini, Michael C. Mozer, Hanie Sedghi, Zachary C. Lipton 외

Recent efforts at explaining the interplay of memorization and generalization in deep overparametrized networks have posited that neural networks $\textit{memorize}$ "hard" examples in the final few layers of the model. …

Memorization