TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
Salient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet.
Code (1)
Tasks
DecoderObjectobject-detectionObject DetectionRGB-D Salient Object DetectionSalient Object DetectionTripletMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Human-Object Interaction Detection via Disentangled Transformer
Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two…
DecoderHuman-Object Interaction DetectionObjectTripletExternal Prompt Features Enhanced Parameter-efficient Fine-tuning for Salient Object Detection
Salient object detection (SOD) aims at finding the most salient objects in images and outputs pixel-level binary masks. Transformer-based methods achieve promising performance due to their global semantic understanding, …
Decoderobject-detectionObject Detectionparameter-efficient fine-tuning+13D Object Discovery and Modeling Using Single RGB-D Images Containing Multiple Object Instances
Unsupervised object modeling is important in robotics, especially for handling a large set of objects. We present a method for unsupervised 3D object discovery, reconstruction, and localization that exploits multiple ins…
DescriptiveObjectObject DiscoveryTripletCoSformer: Detecting Co-Salient Object with Transformers
Co-Salient Object Detection (CoSOD) aims at simulating the human visual system to discover the common and salient objects from a group of relevant images. Recent methods typically develop sophisticated deep learning base…
Co-Salient Object DetectionObjectobject-detectionObject Detection+1Generative Transformer for Accurate and Reliable Salient Object Detection
Transformer, which originates from machine translation, is particularly powerful at modeling long-range dependencies. Currently, the transformer is making revolutionary progress in various vision tasks, leading to signif…
AttributeCamouflaged Object SegmentationGenerative Adversarial NetworkMachine Translation+7