paper-with-me

Papers

Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic Segmentation

2025-07-09 · Joelle Hanna, Damian Borth arxiv

Weakly Supervised Semantic Segmentation (WSSS) is a challenging problem that has been extensively studied in recent years. Traditional approaches often rely on external modules like Class Activation Maps to highlight regions of interest and generate pseudo segmentation masks. In this work, we propose an end-to-end method that directly utilizes the attention maps learned by a Vision Transformer (ViT) for WSSS. We propose training a sparse ViT with multiple [CLS] tokens (one for each class), using a random masking strategy to promote [CLS] token - class assignment. At inference time, we aggregate the different self-attention maps of each [CLS] token corresponding to the predicted labels to generate pseudo segmentation masks. Our proposed approach enhances the interpretability of self-attention maps and ensures accurate class assignments. Extensive experiments on two standard benchmarks and three specialized datasets demonstrate that our method generates accurate pseudo-masks, outperforming related works. Those pseudo-masks can be used to train a segmentation model which achieves results comparable to fully-supervised models, significantly reducing the need for fine-grained labeled data.

📄 PDF Abstract BibTeX arXiv:2507.06848

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

Distract Your Attention: Multi-head Cross Attention Network for Facial Expression Recognition

2021-09-15 · Zhengyao Wen, Wenzhong Lin, Tao Wang, Ge Xu

We present a novel facial expression recognition network, called Distract your Attention Network (DAN). Our method is based on two key observations. Firstly, multiple classes share inherently similar underlying facial ap…

Facial Expression RecognitionFacial Expression Recognition (FER)

Your "Attention" Deserves Attention: A Self-Diversified Multi-Channel Attention for Facial Action Analysis

2022-03-23 · Xiaotian Li, Zhihua Li, Huiyuan Yang, Geran Zhao 외

Visual attention has been extensively studied for learning fine-grained features in both facial expression recognition (FER) and Action Unit (AU) detection. A broad range of previous research has explored how to use atte…

Action AnalysisFacial Expression RecognitionFacial Expression Recognition (FER)

Edit-Your-Interest: Efficient Video Editing via Feature Most-Similar Propagation

2025-10-15 · Yi Zuo, Zitao Wang, Lingling Li, Xu Liu 외 arxiv

Text-to-image (T2I) diffusion models have recently demonstrated significant progress in video editing. However, existing video editing methods are severely limited by their high computational overhead and memory consumpt…

BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way

2024-10-08 · Jiazi Bu, Pengyang Ling, Pan Zhang, Tong Wu 외

The text-to-video (T2V) generation models, offering convenient visual creation, have recently garnered increasing attention. Despite their substantial potential, the generated videos may present artifacts, including stru…

DecoderText-to-Video GenerationVideo Generation

ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way

2025-01-01 · CVPR 2025 1 · Jiazi Bu, Pengyang Ling, Pan Zhang, Tong Wu 외

The text-to-video (T2V) generation models, offering convenient visual creation, have recently garnered increasing attention. Despite their substantial potential, the generated videos may present artifacts, including …

Text-to-Video GenerationVideo Generation