HYDRA -- Hyper Dependency Representation Attentions
Attention is all we need as long as we have enough data. Even so, it is sometimes not easy to determine how much data is enough while the models are becoming larger and larger. In this paper, we propose HYDRA heads, lightweight pretrained linguistic self-attention heads to inject knowledge into transformer models without pretraining them again. Our approach is a balanced paradigm between leaving the models to learn unsupervised and forcing them to conform to linguistic knowledge rigidly as suggested in previous studies. Our experiment proves that the approach is not only the boost performance of the model but also lightweight and architecture friendly. We empirically verify our framework on benchmark datasets to show the contribution of linguistic knowledge to a transformer model. This is a promising result for a new approach to transferring knowledge from linguistic resources into transformer-based models.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Hydra: A method for strain-minimizing hyperbolic embedding of network- and distance-based data
We introduce hydra (hyperbolic distance recovery and approximation), a new method for embedding network- or distance-based data into hyperbolic space. We show mathematically that hydra satisfies a certain optimality guar…
Strain-Minimizing Hyperbolic Network Embeddings with Landmarks
We introduce L-hydra (landmarked hyperbolic distance recovery and approximation), a method for embedding network- or distance-based data into hyperbolic space, which requires only the distance measurements to a few 'land…
HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
The behaviors of deep neural networks (DNNs) are notoriously resistant to human interpretations. In this paper, we propose Hypergradient Data Relevance Analysis, or HYDRA, which interprets the predictions made by DNNs as…
Rolling Shutter CorrectionHydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
Pedestrian analysis plays a vital role in intelligent video surveillance and is a key component for security-centric computer vision systems. Despite that the convolutional neural networks are remarkable in learning disc…
AttributePedestrian Attribute RecognitionPerson Re-IdentificationHydraMamba: Multi-Head State Space Model for Global Point Cloud Learning
The attention mechanism has become a dominant operator in point cloud learning, but its quadratic complexity leads to limited inter-point interactions, hindering long-range dependency modeling between objects. Due to exc…
Long-range modeling