Enriching Language Models with Visually-grounded Word Vectors and the Lancaster Sensorimotor Norms
Language models are trained only on text despite the fact that humans learn their first language in a highly interactive and multimodal environment where the first set of learned words are largely concrete, denoting physical entities and embodied states. To enrich language models with some of this missing experience, we leverage two sources of information: (1) the Lancaster Sensorimotor norms, which provide ratings (means and standard deviations) for over 40,000 English words along several dimensions of embodiment, and which capture the extent to which something is experienced across 11 different sensory modalities, and (2) vectors from coefficients of binary classifiers trained on images for the BERT vocabulary. We pre-trained the ELECTRA model and fine-tuned the RoBERTa model with these two sources of information then evaluate using the established GLUE benchmark and the Visual Dialog benchmark. We find that enriching language models with the Lancaster norms and image vectors improves results in both tasks, with some implications for robust language models that capture holistic linguistic meaning in a language learning context.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual DialogSimilar Papers 제목 키워드 기반
Visual Grounding of Inter-lingual Word-Embeddings
Visual grounding of Language aims at enriching textual representations of language with multiple sources of visual knowledge such as images and videos. Although visual grounding is an area of intense research, inter-ling…
Visual GroundingWord EmbeddingsWord SimilarityCLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
This paper introduces CLIPSwarm, a new algorithm designed to automate the modeling of swarm drone formations based on natural language. The algorithm begins by enriching a provided word, to compose a text prompt that ser…
Word Recognition, Competition, and Activation in a Model of Visually Grounded Speech
In this paper, we study how word-like units are represented and activated in a recurrent neural model of visually grounded speech. The model used in our experiments is trained to project an image and its spoken descripti…
Seeing the advantage: visually grounding word embeddings to better capture human semantic knowledge
Distributional semantic models capture word-level meaning that is useful in many natural language processing tasks and have even been shown to capture cognitive aspects of word meaning. The majority of these models are p…
Grounded language learningImage RetrievalLearning Semantic RepresentationsVisual Grounding+2Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling
This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…
Few-Shot LearningLanguage Acquisition