paper-with-me

홈 › Papers

Self-training solutions for the ICCV 2023 GeoNet Challenge

2023-11-28 · Lijun Sheng, Zhengbo Wang, Jian Liang

GeoNet is a recently proposed domain adaptation benchmark consisting of three challenges (i.e., GeoUniDA, GeoImNet, and GeoPlaces). Each challenge contains images collected from the USA and Asia where there are huge geographical gaps. Our solution adopts a two-stage source-free domain adaptation framework with a Swin Transformer backbone to achieve knowledge transfer from the USA (source) domain to Asia (target) domain. In the first stage, we train a source model using labeled source data with a re-sampling strategy and two types of cross-entropy loss. In the second stage, we generate pseudo labels for unlabeled target data to fine-tune the model. Our method achieves an H-score of 74.56% and ultimately ranks 1st in the GeoUniDA challenge. In GeoImNet and GeoPlaces challenges, our solution also reaches a top-3 accuracy of 64.46% and 51.23%, respectively.

📄 PDF Abstract BibTeX arXiv:2311.16843

Code (1)

tim-learn/geonet23_casia_tim 공식 구현 pytorch

Tasks

Domain AdaptationSource-Free Domain AdaptationTransfer Learning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Stochastic Depth Stochastic Depth aims to shrink the depth of a network during training, while keeping it unchanged during testing. This is achieved by randomly dropping entire…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

Is Self-Supervised Pre-training on Satellite Imagery Better than ImageNet? A Systematic Study with Sentinel-2

2025-02-15 · Saad Lahrichi, Zion Sheng, Shufan Xia, Kyle Bradbury 외

Self-supervised learning (SSL) has demonstrated significant potential in pre-training robust models with limited labeled data, making it particularly valuable for remote sensing (RS) tasks. A common assumption is that pr…

Self-Supervised Learning

GeoNet: Benchmarking Unsupervised Adaptation across Geographies

2023-03-27 · CVPR 2023 1 · Tarun Kalluri, Wangdong Xu, Manmohan Chandraker

In recent years, several efforts have been aimed at improving the robustness of vision models to domains and environments unseen during training. An important practical problem pertains to models deployed in a new geogra…

BenchmarkingDomain Adaptationimage-classificationImage Classification+2

"Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021

2021-10-14 · Ishan Dave, Naman Biyani, Brandon Clark, Rohit Gupta 외

This technical report presents our approach "Knights" to solve the action recognition task on a small subset of Kinetics-400 i.e. Kinetics400ViPriors without using any extra-data. Our approach has 3 main components: stat…

Action RecognitionOptical Flow Estimation

Geometric Representation Learning for Document Image Rectification

2022-10-15 · Hao Feng, Wengang Zhou, Jiajun Deng, Yuechen Wang 외

In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. However, such geometric constraints are largely ignored in existing advanced solutions, which …

Representation Learning

SurgeoNet: Realtime 3D Pose Estimation of Articulated Surgical Instruments from Stereo Images using a Synthetically-trained Network

2024-10-02 · Ahmed Tawfik Aboukhadra, Nadia Robertini, Jameel Malik, Ahmed Elhayek 외

Surgery monitoring in Mixed Reality (MR) environments has recently received substantial focus due to its importance in image-based decisions, skill assessment, and robot-assisted surgery. Tracking hands and articulated s…

3D Pose EstimationMixed RealityPose Estimation