paper-with-me

홈 › Papers

Rethinking FID: Towards a Better Evaluation Metric for Image Generation

2023-11-30 · CVPR 2024 1 · Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, Sanjiv Kumar

As with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of Inception-v3 features of real images, and those of images generated by the algorithm. We highlight important drawbacks of FID: Inception's poor representation of the rich and varied content generated by modern text-to-image models, incorrect normality assumptions, and poor sample complexity. We call for a reevaluation of FID's use as the primary quality metric for generated images. We empirically demonstrate that FID contradicts human raters, it does not reflect gradual improvement of iterative text-to-image models, it does not capture distortion levels, and that it produces inconsistent results when varying the sample size. We also propose an alternative new metric, CMMD, based on richer CLIP embeddings and the maximum mean discrepancy distance with the Gaussian RBF kernel. It is an unbiased estimator that does not make any assumptions on the probability distribution of the embeddings and is sample efficient. Through extensive experiments and analysis, we demonstrate that FID-based evaluations of text-to-image models may be unreliable, and that CMMD offers a more robust and reliable assessment of image quality.

📄 PDF Abstract BibTeX arXiv:2401.09603

Code (4)

google-research/google-research 공식 구현 tf
google-research/google-research/tree/master/cmmd 공식 구현 jax
mazurowski-lab/medical-image-similarity-metrics pytorch
sayakpaul/cmmd-pytorch pytorch

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Average Pooling 설명 없음
Inception-v3 Module Inception-v3 Module is an image block used in the Inception-v3 architecture. This architecture is used on the coarsest (8 ×…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
RBF 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Rethinking and Refining the Distinct Metric

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Distinct is a widely used automatic metric for evaluating the diversity of language generation tasks. However, we observe that the original approach to calculating distinct scores has evident biases that tend to add high…

DiversityText Generation

Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation

2024-09-07 · Jiaxin Cheng, Zixu Zhao, Tong He, Tianjun Xiao 외

Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area with…

Image GenerationLayout-to-Image GenerationVideo Editing

Rethinking HTG Evaluation: Bridging Generation and Recognition

2024-09-04 · Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Marcus Liwicki

The evaluation of generative models for natural image tasks has been extensively studied. Similar protocols and metrics are used in cases with unique particularities, such as Handwriting Generation, even if they might no…

DiversityHandwriting generationHTR

F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration

2024-12-17 · Lu Liu, Huiyu Duan, Qiang Hu, Liu Yang 외

Artificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall…

BenchmarkingFace GenerationImage GenerationImage Quality Assessment

Rethinking the Evaluation of Visible and Infrared Image Fusion

2024-10-09 · Dayan Guan, Yixuan Wu, Tianzhu Liu, Alex C. Kot 외

Visible and Infrared Image Fusion (VIF) has garnered significant interest across a wide range of high-level vision tasks, such as object detection and semantic segmentation. However, the evaluation of VIF methods remains…

object-detectionObject DetectionSegmentationSemantic Segmentation+1