paper-with-me

홈 › Papers

Analyzing the Roles of Language and Vision in Learning from Limited Data

2024-02-15 · Allison Chen, Ilia Sucholutsky, Olga Russakovsky, Thomas L. Griffiths

Does language help make sense of the visual world? How important is it to actually see the world rather than having it described with words? These basic questions about the nature of intelligence have been difficult to answer because we only had one example of an intelligent system -- humans -- and limited access to cases that isolated language or vision. However, the development of sophisticated Vision-Language Models (VLMs) by artificial intelligence researchers offers us new opportunities to explore the contributions that language and vision make to learning about the world. We ablate components from the cognitive architecture of these models to identify their contributions to learning new tasks from limited data. We find that a language model leveraging all components recovers a majority of a VLM's performance, despite its lack of visual input, and that language seems to allow this by providing access to prior knowledge and reasoning.

📄 PDF Abstract BibTeX arXiv:2403.19669

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Visualizing Trends of Key Roles in News Articles

2019-09-12 · IJCNLP 2019 11 · Chen Xia, Haoxiang Zhang, Jacob Moghtader, Allen Wu 외

There are tons of news articles generated every day reflecting the activities of key roles such as people, organizations and political parties. Analyzing these key roles allows us to understand the trends in news. In thi…

Articles

Identifying and Analyzing Task-Encoding Tokens in Large Language Models

2024-01-20 · Yu Bai, Heyan Huang, Cesare Spinoso-Di Piano, Marc-Antoine Rondeau 외

In-context learning (ICL) has become an effective solution for few-shot learning in natural language processing. However, our understanding of ICL's working mechanisms is limited, specifically regarding how models learn …

Computational EfficiencyFew-Shot LearningIn-Context Learning

Video Question Answering with Phrases via Semantic Roles

2021-04-08 · NAACL 2021 4 · Arka Sadhu, Kan Chen, Ram Nevatia

Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA models' application scenario. In this work, w…

Question AnsweringVideo Question Answering

Self-play for Data Efficient Language Acquisition

2020-10-10 · Charles Lovering, Ellie Pavlick

When communicating, people behave consistently across conversational roles: People understand the words they say and are able to produce the words they hear. To date, artificial agents developed for language tasks have l…

Language Acquisition

Open-Vocabulary Argument Role Prediction for Event Extraction

2022-11-03 · Yizhu Jiao, Sha Li, Yiqing Xie, Ming Zhong 외

The argument role in event extraction refers to the relation between an event and an argument participating in it. Despite the great progress in event extraction, existing studies still depend on roles pre-defined by dom…

Event ExtractionLanguage ModellingPrediction