paper-with-me

홈 › Papers

DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning

2025-12-16 · Nakamasa Inoue, Kanoko Goto, Masanari Oi, Martyna Gruszka, Mahiro Ukai, Takumi Hirose, Yusuke Sekikawa arxiv

Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decoder (DISCODE), a novel finetuning-free method that generates robust evaluation scores better aligned with human judgments across diverse domains. The core idea behind DISCODE lies in its test-time adaptive evaluation approach, which introduces the Adaptive Test-Time (ATT) loss, leveraging a Gaussian prior distribution to improve robustness in evaluation score estimation. This loss is efficiently minimized at test time using an analytical solution that we derive. Furthermore, we introduce the Multi-domain Caption Evaluation (MCEval) benchmark, a new image captioning evaluation benchmark covering six distinct domains, designed to assess the robustness of evaluation metrics. In our experiments, we demonstrate that DISCODE achieves state-of-the-art performance as a reference-free evaluation metric across MCEval and four representative existing benchmarks.

📄 PDF Abstract BibTeX arXiv:2512.14420

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

High-Fidelity Music Vocoder using Neural Audio Codecs

2025-02-18 · Luca A. Lanzendörfer, Florian Grötschla, Michael Ungersböck, Roger Wattenhofer

While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages…

DecoderSpeech Synthesis

Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms

2021-05-12 · Ryoto Ishizuka, Ryo Nishikimi, Kazuyoshi Yoshii

This paper describes an automatic drum transcription (ADT) method that directly estimates a tatum-level drum score from a music signal, in contrast to most conventional ADT methods that estimate the frame-level onset pro…

DecoderDrum Transcription

Modelling Context Emotions using Multi-task Learning for Emotion Controlled Dialog Generation

2021-04-01 · EACL 2021 2 · Deeksha Varshney, Asif Ekbal, Pushpak Bhattacharyya

A recent topic of research in natural language generation has been the development of automatic response generation modules that can automatically respond to a user{'}s utterance in an empathetic manner. Previous researc…

DecoderMulti-Task LearningResponse GenerationText Generation

Energy-Aware NECO for Single-Pass Pixel-wise Out-of-Distribution Detection in Semantic Segmentation

2026-05-28 · Boyuan Zhang, Huanshan Huang, Yifei Cao arxiv

Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require …

Out-of-Distribution DetectionSemantic Segmentation

Character-Aware Decoder for Translation into Morphologically Rich Languages

2018-09-06 · WS 2019 8 · Adithya Renduchintala, Pamela Shapiro, Kevin Duh, Philipp Koehn

Neural machine translation (NMT) systems operate primarily on words (or sub-words), ignoring lower-level patterns of morphology. We present a character-aware decoder designed to capture such patterns when translating int…

DecoderMachine TranslationNMTTranslation