paper-with-me

홈 › Papers

CoVLM: Leveraging Consensus from Vision-Language Models for Semi-supervised Multi-modal Fake News Detection

2024-10-06 · Devank, Jayateja Kalla, Soma Biswas

In this work, we address the real-world, challenging task of out-of-context misinformation detection, where a real image is paired with an incorrect caption for creating fake news. Existing approaches for this task assume the availability of large amounts of labeled data, which is often impractical in real-world, since it requires extensive manual intervention and domain expertise. In contrast, since obtaining a large corpus of unlabeled image-text pairs is much easier, here, we propose a semi-supervised protocol, where the model has access to a limited number of labeled image-text pairs and a large corpus of unlabeled pairs. Additionally, the occurrence of fake news being much lesser compared to the real ones, the datasets tend to be highly imbalanced, thus making the task even more challenging. Towards this goal, we propose a novel framework, Consensus from Vision-Language Models (CoVLM), which generates robust pseudo-labels for unlabeled pairs using thresholds derived from the labeled data. This approach can automatically determine the right threshold parameters of the model for selecting the confident pseudo-labels. Experimental results on benchmark datasets across challenging conditions and comparisons with state-of-the-art approaches demonstrate the effectiveness of our framework.

📄 PDF Abstract BibTeX arXiv:2410.04426

Code (1)

devank3/CoVLM 공식 구현

Tasks

Fake News DetectionMisinformation

Similar Papers 제목 키워드 기반

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

2023-11-06 · Junyan Li, Delin Chen, Yining Hong, Zhenfang Chen 외

A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundation models (VLMs) fall short of such compo…

CoLAQuestion AnsweringReferring ExpressionReferring Expression Comprehension+3

COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning

2025-12-10 · Lin Li, Yuxin Cai, Jianwu Fang, Jianru Xue 외 arxiv

End-to-end autonomous driving frameworks face persistent challenges in generalization, training efficiency, and interpretability. While recent methods leverage Vision-Language Models (VLMs) through supervised learning on…

Reinforcement LearningAutonomous Driving

VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Human Annotation-Free Pathological Image Classification

2024-03-23 · Lanfeng Zhong, Xin Liao, Shaoting Zhang, Xiaofan Zhang 외

Despite that deep learning methods have achieved remarkable performance in pathology image classification, they heavily rely on labeled data, demanding extensive human annotation efforts. In this study, we present a nove…

image-classificationImage Classificationzero-shot-classificationZero-Shot Learning

Co-Training Vision Language Models for Remote Sensing Multi-task Learning

2025-11-26 · Qingyun Li, Shuran Ma, Junwei Luo, Yi Yu 외 arxiv

With Transformers achieving outstanding performance on individual remote sensing (RS) tasks, we are now approaching the realization of a unified model that excels across multiple tasks through multi-task learning (MTL). …

Multi-Task LearningObject Detection

SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

2026-05-26 · Yuqi Liu, Yufei Chen, Wei Fu, Xiaodong Yue 외 arxiv

Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi-Supervised Learning (SSL). However, due to significant inter-sample morphological variability, existing S…

Pancreas Segmentation