paper-with-me

Papers

Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption Similarity

2022-10-03 · Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu

Automatic Audio Captioning (AAC) refers to the task of translating an audio sample into a natural language (NL) text that describes the audio events, source of the events and their relationships. Unlike NL text generation tasks, which rely on metrics like BLEU, ROUGE, METEOR based on lexical semantics for evaluation, the AAC evaluation metric requires an ability to map NL text (phrases) that correspond to similar sounds in addition lexical semantics. Current metrics used for evaluation of AAC tasks lack an understanding of the perceived properties of sound represented by text. In this paper, wepropose a novel metric based on Text-to-Audio Grounding (TAG), which is, useful for evaluating cross modal tasks like AAC. Experiments on publicly available AAC data-set shows our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature.

📄 PDF Abstract BibTeX arXiv:2210.06354

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningImage CaptioningTAGText Generation

Similar Papers 제목 키워드 기반

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

2026-06-17 · Sakshi Joshi, Dhruv Subhash Rathi, Sanskar Singh, Eldho Ittan George 외 arxiv

AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowle…

Speech Recognition

ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability

2026-03-20 · Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien, Toshiki Ogimoto 외 arxiv

While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. …

Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamics Audio-Visual Scenarios

2023-05-21 · Yuanyuan Jiang, Jianqin Yin

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for s…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringScene Understanding+1

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

2025-07-01 · Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat 외 arxiv

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…

Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual Grounding

Audio-visual training for improved grounding in video-text LLMs

2024-07-21 · Shivprasad Sagare, Hemachandran S, Kinshuk Sarabhai, Prashant Ullegaddi 외

Recent advances in multimodal LLMs, have led to several video-text models being proposed for critical video-related tasks. However, most of the previous works support visual input only, essentially muting the audio signa…

Video Understanding