paper-with-me

Papers

A Visual Representation-guided Framework with Global Affinity for Weakly Supervised Salient Object Detection

2023-02-21 · Binwei Xu, Haoran Liang, Weihua Gong, Ronghua Liang, Peng Chen

Fully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance, scribble-based SOD methods have attracted increasing attention. Previous scribble-based models directly implement the SOD task only based on SOD training data with limited information, it is extremely difficult for them to understand the image and further achieve a superior SOD task. In this paper, we propose a simple yet effective framework guided by general visual representations with rich contextual semantic knowledge for scribble-based SOD. These general visual representations are generated by self-supervised learning based on large-scale unlabeled datasets. Our framework consists of a task-related encoder, a general visual module, and an information integration module to efficiently combine the general visual representations with task-related features to perform the SOD task based on understanding the contextual connections of images. Meanwhile, we propose a novel global semantic affinity loss to guide the model to perceive the global structure of the salient objects. Experimental results on five public benchmark datasets demonstrate that our method, which only utilizes scribble annotations without introducing any extra label, outperforms the state-of-the-art weakly supervised SOD methods. Specifically, it outperforms the previous best scribble-based method on all datasets with an average gain of 5.5% for max f-measure, 5.8% for mean f-measure, 24% for MAE, and 3.1% for E-measure. Moreover, our method achieves comparable or even superior performance to the state-of-the-art fully supervised models.

📄 PDF Abstract BibTeX arXiv:2302.10697

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSalient Object DetectionSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

MAE 설명 없음

Similar Papers 제목 키워드 기반

Curvature-Guided Geometric Representation for Protein-Ligand Binding Affinity Prediction

2026-06-12 · Shuai Li, Chuan-Xian Ren, Yuhao Li, Ziqi Huang 외 arxiv

Protein-ligand binding affinity (PLA) prediction is critical in drug discovery. Despite the notable advancements in machine learning-based approaches, existing methods struggle to jointly characterize local geometric org…

Drug Discovery

Deep Class-Specific Affinity-Guided Convolutional Network for Multimodal Unpaired Image Segmentation

2021-01-05 · Jingkun Chen, Wenqi Li, Hongwei Li, JianGuo Zhang

Multi-modal medical image segmentation plays an essential role in clinical diagnosis. It remains challenging as the input modalities are often not well-aligned spatially. Existing learning-based methods mainly consider s…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Learning Protein-Ligand Binding in Hyperbolic Space

2025-08-21 · Jianhui Wang, Wenyu Zhu, Bowen Gao, Xin Hong 외 arxiv

Protein-ligand binding prediction is central to virtual screening and affinity ranking, two fundamental tasks in drug discovery. While recent retrieval-based methods embed ligands and protein pockets into Euclidean space…

Representation LearningDrug Discovery

Natural Image Matting via Guided Contextual Attention

2020-01-13 · Yaoyi Li, Hongtao Lu

Over the last few years, deep learning based approaches have achieved outstanding improvements in natural image matting. Many of these methods can generate visually plausible alpha estimations, but typically yield blurry…

Image MattingSemantic Image MattingTransparent objects

Looking into Your Speech: Learning Cross-modal Affinity for Audio-visual Speech Separation

2021-03-25 · CVPR 2021 1 · Jiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang 외

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared informat…

Audio-Visual SynchronizationSpeech Separation