paper-with-me

Papers

Accommodating Audio Modality in CLIP for Multimodal Processing

2023-03-12 · Ludan Ruan, Anwen Hu, Yuqing Song, Liang Zhang, Sipeng Zheng, Qin Jin

Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training, as introducing more modalities can greatly complicate model design and optimization. In this paper, we extend the stateof-the-art Vision-Language model CLIP to accommodate the audio modality for Vision-Language-Audio multimodal processing. Specifically, we apply inter-modal and intra-modal contrastive learning to explore the correlation between audio and other modalities in addition to the inner characteristics of the audio modality. Moreover, we further design an audio type token to dynamically learn different audio information type for different scenarios, as both verbal and nonverbal heterogeneous information is conveyed in general audios. Our proposed CLIP4VLA model is validated in different downstream tasks including video retrieval and video captioning, and achieves the state-of-the-art performance on the benchmark datasets of MSR-VTT, VATEX, and Audiocaps.

📄 PDF Abstract BibTeX arXiv:2303.06591

Code (1)

ludanruan/clip4vla 공식 구현 pytorch

Tasks

AudioCapsContrastive LearningLanguage ModelingLanguage ModellingRetrievalVideo CaptioningVideo Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

2026-05-18 · Morunliu Yang, Ruotao Xu, Le Li, Yue Wang 외 arxiv

Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhea…

What Are They Doing? Joint Audio-Speech Co-Reasoning

2024-09-22 · Yingzhi Wang, Pooneh Mousavi, Artem Ploujnikov, Mirco Ravanelli

In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have ma…

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2024-12-19 · CVPR 2025 1 · Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya 외

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…

Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound Generation

Multimodality and Attention Increase Alignment in Natural Language Prediction Between Humans and Computational Models

2023-08-11 · Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards 외

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are…

Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion

2024-05-09 · Syed Hammad Ahmed, Muhammad Junaid Khan, Gita Sukthankar

Due to the rise in video content creation targeted towards children, there is a need for robust content moderation schemes for video hosting platforms. A video that is visually benign may include audio content that is in…

Prompt LearningRobust classification