paper-with-me

Papers

Large Language Models Meet Computer Vision: A Brief Survey

2023-11-28 · Raby Hamadi

Recently, the intersection of Large Language Models (LLMs) and Computer Vision (CV) has emerged as a pivotal area of research, driving significant advancements in the field of Artificial Intelligence (AI). As transformers have become the backbone of many state-of-the-art models in both Natural Language Processing (NLP) and CV, understanding their evolution and potential enhancements is crucial. This survey paper delves into the latest progressions in the domain of transformers and their subsequent successors, emphasizing their potential to revolutionize Vision Transformers (ViTs) and LLMs. This survey also presents a comparative analysis, juxtaposing the performance metrics of several leading paid and open-source LLMs, shedding light on their strengths and areas of improvement as well as a literature review on how LLMs are being used to tackle vision related tasks. Furthermore, the survey presents a comprehensive collection of datasets employed to train LLMs, offering insights into the diverse data available to achieve high performance in various pre-training and downstream tasks of LLMs. The survey is concluded by highlighting open directions in the field, suggesting potential venues for future research and development. This survey aims to underscores the profound intersection of LLMs on CV, leading to a new era of integrated and advanced AI models.

📄 PDF Abstract BibTeX arXiv:2311.16673

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

Graph Meets LLMs: Towards Large Graph Models

2023-08-28 · Ziwei Zhang, Haoyang Li, Zeyang Zhang, Yijian Qin 외

Large models have emerged as the most recent groundbreaking achievements in artificial intelligence, and particularly machine learning. However, when it comes to graphs, large models have not achieved the same level of s…

Using Computer Vision to Analyze Non-manual Marking of Questions in KRSL

2021-08-01 · MTSummit 2021 8 · Anna Kuznetsova, Alfarabi Imashev, Medet Mukushev, Anara Sandygulova 외

This paper presents a study that compares non-manual markers of polar and wh-questions to statements in Kazakh-Russian Sign Language (KRSL) in a dataset collected for NLP tasks. The primary focus of the study is to demon…

MeetUp! A Corpus of Joint Activity Dialogues in a Visual Environment

2019-07-11 · Nikolai Ilinykh, Sina Zarrieß, David Schlangen

Building computer systems that can converse about their visual environment is one of the oldest concerns of research in Artificial Intelligence and Computational Linguistics (see, for example, Winograd's 1972 SHRDLU syst…

Vision-and-Language Pretrained Models: A Survey

2022-04-15 · Siqu Long, Feiqi Cao, Soyeon Caren Han, Haiqin Yang

Pretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision and language pretraining by feeding visual…

Survey

Challenges in Designing Natural Language Interfaces for Complex Visual Models

2021-04-01 · EACL (HCINLP) 2021 4 · Henrik Voigt, Monique Meuschke, Kai Lawonn, Sina Zarrieß

Intuitive interaction with visual models becomes an increasingly important task in the field of Visualization (VIS) and verbal interaction represents a significant aspect of it. Vice versa, modeling verbal interaction in…