paper-with-me

홈 › Papers

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

2024-12-13 · Sara Ghazanfari, Siddharth Garg, Nicolas Flammarion, Prashanth Krishnamurthy, Farshad Khorrami, Francesco Croce

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal models (LMMs), can be applied as zero-shot perceptual metrics, and several recent works have developed models specialized in narrow perceptual tasks. However, the extent to which existing perceptual metrics align with human perception remains unclear. To investigate this question, we introduce UniSim-Bench, a benchmark encompassing 7 multi-modal perceptual similarity tasks, with a total of 25 datasets. Our evaluation reveals that while general-purpose models perform reasonably well on average, they often lag behind specialized models on individual tasks. Conversely, metrics fine-tuned for specific tasks fail to generalize well to unseen, though related, tasks. As a first step towards a unified multi-task perceptual similarity metric, we fine-tune both encoder-based and generative vision-language models on a subset of the UniSim-Bench tasks. This approach yields the highest average performance, and in some cases, even surpasses taskspecific models. Nevertheless, these models still struggle with generalization to unseen tasks, highlighting the ongoing challenge of learning a robust, unified perceptual similarity metric capable of capturing the human notion of similarity. The code and models are available at https://github.com/SaraGhazanfari/UniSim.

📄 PDF Abstract BibTeX arXiv:2412.10594

Code (1)

saraghazanfari/unisim 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

2025-05-06 · Kapil Wanaskar, Gaytri Jena, Magdalini Eirinaki

This work presents an open-source unified benchmarking and evaluation framework for text-to-image generation models, with a particular focus on the impact of metadata augmented prompts. Leveraging the DeepFashion-MultiMo…

BenchmarkingImage GenerationMLLM Aesthetic EvaluationModel Selection+4

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

2026-05-03 · Xiaoda Yang, Majun Zhang, Changhao Pan, Nick Huang 외 arxiv

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-da…

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

2025-12-25 · Shuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image f…

Visual Question AnsweringText-to-Image GenerationVisual Grounding

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

2025-12-24 · Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 외 arxiv

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or na…

Instruction FollowingVideo Generation

MedSR-Vision: Deep Learning Framework for Multi-Domain Medical Image Super-Resolution

2026-05-05 · Subhash Gurappa, Trivikram Satharasi, Yashas Hariprasad, Sundararaj Sitharama Iyengar arxiv

Medical image super-resolution (MedSR) is essential for improving diagnostic precision across diverse imaging modalities such as MRI, CT, X-ray, Ultrasound, and Fundus imaging. Despite rapid advances in deep learning, ch…

Image Super-Resolution