paper-with-me

홈 › Papers

Learning Tree-Structured Composition of Data Augmentation

2024-08-26 · Dongyue Li, Kailai Chen, Predrag Radivojac, Hongyang R. Zhang

Data augmentation is widely used for training a neural network given little labeled data. A common practice of augmentation training is applying a composition of multiple transformations sequentially to the data. Existing augmentation methods such as RandAugment randomly sample from a list of pre-selected transformations, while methods such as AutoAugment apply advanced search to optimize over an augmentation set of size $k^d$, which is the number of transformation sequences of length $d$, given a list of $k$ transformations. In this paper, we design efficient algorithms whose running time complexity is much faster than the worst-case complexity of $O(k^d)$, provably. We propose a new algorithm to search for a binary tree-structured composition of $k$ transformations, where each tree node corresponds to one transformation. The binary tree generalizes sequential augmentations, such as the SimCLR augmentation scheme for contrastive learning. Using a top-down, recursive search procedure, our algorithm achieves a runtime complexity of $O(2^d k)$, which is much faster than $O(k^d)$ as $k$ increases above $2$. We apply our algorithm to tackle data distributions with heterogeneous subpopulations by searching for one tree in each subpopulation and then learning a weighted combination, resulting in a forest of trees. We validate our proposed algorithms on numerous graph and image datasets, including a multi-label graph classification dataset we collected. The dataset exhibits significant variations in the sizes of graphs and their average degrees, making it ideal for studying data augmentation. We show that our approach can reduce the computation cost by 43% over existing search methods while improving performance by 4.3%. The tree structures can be used to interpret the relative importance of each transformation, such as identifying the important transformations on small vs. large graphs.

📄 PDF Abstract BibTeX arXiv:2408.14381

Code (1)

virtuosoresearch/tree-data-augmentation 공식 구현 pytorch

Tasks

Contrastive LearningData AugmentationGraph Classification

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Bitcoin Customer Service Number +1-833-534-1729 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Tanh Activation 설명 없음
Kaiming Initialization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

TreeMix: Compositional Constituency-based Data Augmentation for Natural Language Understanding

2022-05-12 · NAACL 2022 7 · Le Zhang, Zichao Yang, Diyi Yang

Data augmentation is an effective approach to tackle over-fitting. Many previous works have proposed different data augmentations strategies for NLP, such as noise injection, word replacement, back-translation etc. Thoug…

Constituency ParsingData AugmentationDiversityNatural Language Understanding+2

SUBS: Subtree Substitution for Compositional Semantic Parsing

2022-05-03 · NAACL 2022 7 · Jingfeng Yang, Le Zhang, Diyi Yang

Although sequence-to-sequence models often achieve good performance in semantic parsing for i.i.d. data, their performance is still inferior in compositional generalization. Several data augmentation methods have been pr…

Data AugmentationSemantic Parsing

SUBS: Subtree Substitution for Compositional Semantic Parsing

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Although sequence-to-sequence models often achieve good performance in semantic parsing for i.i.d. data, their performance is still inferior in compositional generalization. Several data augmentation methods have been pr…

Data AugmentationSemantic Parsing

Tree-structured composition in neural networks without tree-structured architectures

2015-06-16 · Samuel R. Bowman, Christopher D. Manning, Christopher Potts

Tree-structured neural networks encode a particular tree geometry for a sentence in the network design. However, these models have at best only slightly outperformed simpler sequence-based models. We hypothesize that neu…

Sentence

Tensor Decompositions in Recursive Neural Networks for Tree-Structured Data

2020-06-18 · Daniele Castellana, Davide Bacciu

The paper introduces two new aggregation functions to encode structural knowledge from tree-structured data. They leverage the Canonical and Tensor-Train decompositions to yield expressive context aggregation while limit…

General Classification