Exploiting Domain Knowledge via Grouped Weight Sharing with Application to Text Categorization
A fundamental advantage of neural models for NLP is their ability to learn representations from scratch. However, in practice this often means ignoring existing external linguistic resources, e.g., WordNet or domain specific ontologies such as the Unified Medical Language System (UMLS). We propose a general, novel method for exploiting such resources via weight sharing. Prior work on weight sharing in neural networks has considered it largely as a means of model compression. In contrast, we treat weight sharing as a flexible mechanism for incorporating prior knowledge into neural models. We show that this approach consistently yields improved performance on classification tasks compared to baseline strategies that do not exploit weight sharing.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationModel CompressionText CategorizationSimilar Papers 제목 키워드 기반
Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition
While transformers and their variant conformers show promising performance in speech recognition, the parameterized property leads to much memory cost during training and inference. Some works use cross-layer weight-shar…
Knowledge DistillationMixture-of-Expertsspeech-recognitionSpeech RecognitionExploiting Both Domain-specific and Invariant Knowledge via a Win-win Transformer for Unsupervised Domain Adaptation
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Most existing UDA approaches enable knowledge transfer via learning domain-invariant representat…
Domain AdaptationTransfer LearningUnsupervised Domain AdaptationData Similarity-Based One-Shot Clustering for Multi-Task Hierarchical Federated Learning
We address the problem of cluster identity estimation in a hierarchical federated learning setting in which users work toward learning different tasks. To overcome the challenge of task heterogeneity, users need to be gr…
ClusteringFederated LearningKnowledge Sharing via Social Login: Exploiting Microblogging Service for Warming up Social Question Answering Websites
Grouped Adaptive Loss Weighting for Person Search
Person search is an integrated task of multiple sub-tasks such as foreground/background classification, bounding box regression and person re-identification. Therefore, person search is a typical multi-task learning prob…
Model OptimizationMulti-Task LearningPerson Re-IdentificationPerson Search