paper-with-me

Papers

Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models

2023-03-28 · Adyasha Maharana, Amita Kamath, Christopher Clark, Mohit Bansal, Aniruddha Kembhavi

As general purpose vision models get increasingly effective at a wide set of tasks, it is imperative that they be consistent across the tasks they support. Inconsistent AI models are considered brittle and untrustworthy by human users and are more challenging to incorporate into larger systems that take dependencies on their outputs. Measuring consistency between very heterogeneous tasks that might include outputs in different modalities is challenging since it is difficult to determine if the predictions are consistent with one another. As a solution, we introduce a benchmark dataset, CocoCon, where we create contrast sets by modifying test instances for multiple tasks in small but semantically meaningful ways to change the gold label and outline metrics for measuring if a model is consistent by ranking the original and perturbed instances across tasks. We find that state-of-the-art vision-language models suffer from a surprisingly high degree of inconsistent behavior across tasks, especially for more heterogeneous tasks. To alleviate this issue, we propose a rank correlation-based auxiliary training objective, computed over large automatically created cross-task contrast sets, that improves the multi-task consistency of large unified models while retaining their original accuracy on downstream tasks.

📄 PDF Abstract BibTeX arXiv:2303.16133

Code (1)

adymaharana/cococon 공식 구현 jax

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Multi-view Graph Learning by Joint Modeling of Consistency and Inconsistency

2020-08-24 · Youwei Liang, Dong Huang, Chang-Dong Wang, Philip S. Yu

Graph learning has emerged as a promising technique for multi-view clustering with its ability to learn a unified and robust graph from multiple views. However, existing graph learning methods mostly focus on the multi-v…

ClusteringGraph Learning

Inconsistency Matters: A Knowledge-guided Dual-inconsistency Network for Multi-modal Rumor Detection

2021-11-01 · Findings (EMNLP) 2021 11 · Mengzhu Sun, Xi Zhang, Jianqiang Ma, Yazheng Liu

Rumor spreaders are increasingly utilizing multimedia content to attract the attention and trust of news consumers. Though a set of rumor detection models have exploited the multi-modal data, they seldom consider the inc…

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

2026-06-03 · Zixu Li, Yupeng Hu, Zhiwei Chen, Haokun Wen 외 arxiv

Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases whe…

Image Retrieval

Exposing Text-Image Inconsistency Using Diffusion Models

2024-04-28 · Mingzhen Huang, Shan Jia, Zhou Zhou, Yan Ju 외

In the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based metho…

Misinformation

FairAdapter: Detecting AI-generated Images with Improved Fairness

2024-11-22 · Feng Ding, Jun Zhang, Xinan He, Jianfeng Xu

The high-quality, realistic images generated by generative models pose significant challenges for exposing them.So far, data-driven deep neural networks have been justified as the most efficient forensics tools for the c…

Fairness