paper-with-me

Papers

FLAP: Fast Language-Audio Pre-training

2023-11-02 · Ching-Feng Yeh, Po-Yao Huang, Vasu Sharma, Shang-Wen Li, Gargi Gosh

We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, contrastive learning and reconstruction. For efficiency, FLAP randomly drops audio spectrogram tokens, focusing solely on the remaining ones for self-supervision. Through inter-modal contrastive learning, FLAP learns to align paired audio and text representations in a shared latent space. Notably, FLAP leverages multiple augmented views via masking for inter-modal contrast and learns to reconstruct the masked portion of audio tokens. Moreover, FLAP leverages large language models (LLMs) to augment the text inputs, contributing to improved performance. These approaches lead to more robust and informative audio-text representations, enabling FLAP to achieve state-of-the-art (SoTA) performance on audio-text retrieval tasks on AudioCaps (achieving 53.0% R@1) and Clotho (achieving 25.5% R@1).

📄 PDF Abstract BibTeX arXiv:2311.01615

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsContrastive LearningRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Linear Representation Meta-Reinforcement Learning for Instant Adaptation

2021-01-12 · Matt Peng, Banghua Zhu, Jiantao Jiao

This paper introduces Fast Linearized Adaptive Policy (FLAP), a new meta-reinforcement learning (meta-RL) method that is able to extrapolate well to out-of-distribution tasks without the need to reuse data from training,…

continuous-controlContinuous ControlMeta Reinforcement Learningreinforcement-learning+2

Fluctuation-based Adaptive Structured Pruning for Large Language Models

2023-12-19 · Yongqi An, Xu Zhao, Tao Yu, Ming Tang 외

Network Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost a…

Network Pruning

Evolution Strategies for Deep RL pretraining

2026-03-31 · Adrian Martínez, Ananya Gupta, Hanka Goralija, Mario Rico 외 arxiv

Although Deep Reinforcement Learning has proven highly effective for complex decision-making problems, it demands significant computational resources and careful parameter adjustment in order to develop successful strate…

Reinforcement Learning

Data-Driven Approaches for Thrust Prediction in Underwater Flapping Fin Propulsion Systems

2024-06-04 · Julian Lee, Kamal Viswanath, Alisha Sharma, Jason Geder 외

Flapping-fin underwater vehicle propulsion systems provide an alternative to propeller-driven systems in situations that require involve a constrained environment or require high maneuverability. Testing new configuratio…

Wear Classification of Abrasive Flap Wheels using a Hierarchical Deep Learning Approach

2026-03-13 · Falko Kähler, Maxim Wille, Ole Schmedemann, Thorsten Schüppstuhl arxiv

Abrasive flap wheels are common for finishing complex free-form surfaces due to their flexibility. However, this flexibility results in complex wear patterns such as concave/convex flap profiles or flap tears, which infl…

Transfer Learning