paper-with-me

홈 › Papers

Visual Atoms: Pre-training Vision Transformers with Sinusoidal Waves

2023-03-02 · CVPR 2023 1 · Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka, Rio Yokota

Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision transformers. However, the lack of a systematic investigation as to why these contour-oriented synthetic datasets can achieve the same accuracy as real datasets leaves much room for skepticism. In the present work, we develop a novel methodology based on circular harmonics for systematically investigating the design space of contour-oriented synthetic datasets. This allows us to efficiently search the optimal range of FDSL parameters and maximize the variety of synthetic images in the dataset, which we found to be a critical factor. When the resulting new dataset VisualAtom-21k is used for pre-training ViT-Base, the top-1 accuracy reached 83.7% when fine-tuning on ImageNet-1k. This is close to the top-1 accuracy (84.2%) achieved by JFT-300M pre-training, while the number of images is 1/14. Unlike JFT-300M which is a static dataset, the quality of synthetic datasets will continue to improve, and the current work is a testament to this possibility. FDSL is also free of the common issues associated with real images, e.g. privacy/copyright issues, labeling costs/errors, and ethical biases.

📄 PDF Abstract BibTeX arXiv:2303.01112

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sinusoidal Initialization, Time for a New Start

2025-05-19 · Alberto Fernández-Hernández, Jose I. Mestre, Manuel F. Dolz, Jose Duato 외

Initialization plays a critical role in Deep Neural Network training, directly influencing convergence, stability, and generalization. Common approaches such as Glorot and He initializations rely on randomness, which can…

Explainability Techniques for Chemical Language Models

2023-05-25 · Stefan Hödl, William Robinson, Yoram Bachrach, Wilhelm Huck 외

Explainability techniques are crucial in gaining insights into the reasons behind the predictions of deep learning models, which have not yet been applied to chemical language models. We propose an explainable AI techniq…

ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models

2025-02-10 · Ehsan Zeraatkar, Salah Faroughi, Jelena Tešić

Purpose: Earth system models (ESMs) integrate the interactions of the atmosphere, ocean, land, ice, and biosphere to estimate the state of regional and global climate under a wide variety of conditions. The ESMs are high…

Image ReconstructionSSIM

Sparse Fine-Tuning of Transformers for Generative Tasks

2025-07-14 · Wei Chen, Jingxi Yu, Zichen Miao, Qiang Qiu arxiv

Large pre-trained transformers have revolutionized artificial intelligence across various domains, and fine-tuning remains the dominant approach for adapting these models to downstream tasks due to the cost of training f…

Image Editing

Pattern Encoding on the Poincare Sphere

2014-10-01 · Aleksandra Pizurica

This paper presents a convenient graphical tool for encoding visual patterns (such as image patches and image atoms) as point constellations in a space spanned by perceptual features and with a clear geometrical interpre…

Clustering