paper-with-me

홈 › Papers

Discovering Spatial Relationships by Transformers for Domain Generalization

2021-08-23 · Cuicui Kang, Karthik Nandakumar

Due to the rapid increase in the diversity of image data, the problem of domain generalization has received increased attention recently. While domain generalization is a challenging problem, it has achieved great development thanks to the fast development of AI techniques in computer vision. Most of these advanced algorithms are proposed with deep architectures based on convolution neural nets (CNN). However, though CNNs have a strong ability to find the discriminative features, they do a poor job of modeling the relations between different locations in the image due to the response to CNN filters are mostly local. Since these local and global spatial relationships are characterized to distinguish an object under consideration, they play a critical role in improving the generalization ability against the domain gap. In order to get the object parts relationships to gain better domain generalization, this work proposes to use the self attention model. However, the attention models are proposed for sequence, which are not expert in discriminate feature extraction for 2D images. Considering this, we proposed a hybrid architecture to discover the spatial relationships between these local features, and derive a composite representation that encodes both the discriminative features and their relationships to improve the domain generalization. Evaluation on three well-known benchmarks demonstrates the benefits of modeling relationships between the features of an image using the proposed method and achieves state-of-the-art domain generalization performance. More specifically, the proposed algorithm outperforms the state-of-the-art by 2.2% and 3.4% on PACS and Office-Home databases, respectively.

📄 PDF Abstract BibTeX arXiv:2108.10046

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Discovering Generalizable Spatial Goal Representations via Graph-based Active Reward Learning

2022-11-24 · Aviv Netanyahu, Tianmin Shu, Joshua Tenenbaum, Pulkit Agrawal

In this work, we consider one-shot imitation learning for object rearrangement tasks, where an AI agent needs to watch a single expert demonstration and learn to perform the same task in different environments. To achiev…

AI AgentImitation LearningObject Rearrangement

Deep Spatial Domain Generalization

2022-10-03 · Dazhou Yu, Guangji Bai, Yun Li, Liang Zhao

Spatial autocorrelation and spatial heterogeneity widely exist in spatial data, which make the traditional machine learning model perform badly. Spatial domain generalization is a spatial extension of domain generalizati…

Domain GeneralizationGraph Neural NetworkSpatial Interpolation

A Video Is Worth Three Views: Trigeminal Transformers for Video-based Person Re-identification

2021-04-05 · Xuehu Liu, Pingping Zhang, Chenyang Yu, Huchuan Lu 외

Video-based person re-identification (Re-ID) aims to retrieve video sequences of the same person under non-overlapping cameras. Previous methods usually focus on limited views, such as spatial, temporal or spatial-tempor…

Person Re-IdentificationVideo-Based Person Re-Identification

Invenio: Discovering Hidden Relationships Between Tasks/Domains Using Structured Meta Learning

2019-11-24 · Sameeksha Katoch, Kowshik Thopalli, Jayaraman J. Thiagarajan, Pavan Turaga 외

Exploiting known semantic relationships between fine-grained tasks is critical to the success of recent model agnostic approaches. These approaches often rely on meta-optimization to make a model robust to systematic tas…

Few-Shot LearningMeta-LearningSelf-Supervised Learning

Cameras as Relative Positional Encoding

2025-07-14 · RuiLong Li, Brent Yi, Junchen Liu, Hang Gao 외

Transformers are increasingly prevalent for multi-view computer vision tasks, where geometric relationships between viewpoints are critical for 3D perception. To leverage these relationships, multi-view transformers must…

Depth EstimationNovel View SynthesisStereo Depth Estimation