paper-with-me

홈 › Papers

Teaching Structured Vision&Language Concepts to Vision&Language Models

2022-11-21 · Sivan Doveh, Assaf Arbelle, Sivan Harary, Rameswar Panda, Roei Herzig, Eli Schwartz, Donghyun Kim, Raja Giryes, Rogerio Feris, Shimon Ullman, Leonid Karlinsky

Vision and Language (VL) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective notion of Structured Vision&Language Concepts (SVLC) which includes object attributes, relations, and states which are present in the text and visible in the image. Recent studies have shown that even the best VL models struggle with SVLC. A possible way of fixing this issue is by collecting dedicated datasets for teaching each SVLC type, yet this might be expensive and time-consuming. Instead, we propose a more elegant data-driven approach for enhancing VL models' understanding of SVLCs that makes more effective use of existing VL pre-training datasets and does not require any additional data. While automatic understanding of image structure still remains largely unsolved, language structure is much better modeled and understood, allowing for its effective utilization in teaching VL models. In this paper, we propose various techniques based on language structure understanding that can be used to manipulate the textual part of off-the-shelf paired VL datasets. VL models trained with the updated data exhibit a significant improvement of up to 15% in their SVLC understanding with only a mild degradation in their zero-shot capabilities both when training from scratch or fine-tuning a pre-trained model.

📄 PDF Abstract BibTeX arXiv:2211.11733

Code (1)

sivandoveh/tsvlc 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Teaching Structured Vision & Language Concepts to Vision & Language Models

2023-01-01 · CVPR 2023 1 · Sivan Doveh, Assaf Arbelle, Sivan Harary, Eli Schwartz 외

Vision and Language (VL) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective …

Automatic Teaching Platform on Vision Language Retrieval Augmented Generation

2025-03-07 · Ruslan Gokhman, Jialu Li, Youshan Zhang

Automating teaching presents unique challenges, as replicating human interaction and adaptability is complex. Automated systems cannot often provide nuanced, real-time feedback that aligns with students' individual learn…

RAGRetrievalRetrieval-augmented Generation

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

2025-05-14 · Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti 외

Large language models have become multimodal, and many of them are said to integrate their modalities using common representations. If this were true, a drawing of a car as an image, for instance, should map to the simil…

ConStruct-VL: Data-Free Continual Structured VL Concepts Learning

2022-11-17 · CVPR 2023 1 · James Seale Smith, Paola Cascante-Bonilla, Assaf Arbelle, Donghyun Kim 외

Recently, large-scale pre-trained Vision-and-Language (VL) foundation models have demonstrated remarkable capabilities in many zero-shot downstream tasks, achieving competitive results for recognizing objects defined by …

TEACHING -- Trustworthy autonomous cyber-physical applications through human-centred intelligence

2021-07-14 · Davide Bacciu, Siranush Akarmazyan, Eric Armengaud, Manlio Bacco 외

This paper discusses the perspective of the H2020 TEACHING project on the next generation of autonomous applications running in a distributed and highly heterogeneous environment comprising both virtual and physical reso…

Federated Learning