paper-with-me

Papers

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

2022-10-17 · Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao

This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding; ($ii$) VLP for core computer vision tasks, such as (open-set) image classification, object detection, and segmentation; and ($iii$) VLP for video-text tasks, such as video captioning, video-text retrieval, and video question answering. For each category, we present a comprehensive review of state-of-the-art methods, and discuss the progress that has been made and challenges still being faced, using specific systems and models as case studies. In addition, for each category, we discuss advanced topics being actively explored in the research community, such as big foundation models, unified modeling, in-context few-shot learning, knowledge, robustness, and computer vision in the wild, to name a few.

📄 PDF Abstract BibTeX arXiv:2210.09263

Code (1)

computer-vision-in-the-wild/cvinw_readings 공식 구현

Tasks

Few-Shot LearningImage Captioningimage-classificationImage ClassificationImage-text Retrievalobject-detectionObject DetectionQuestion AnsweringRetrievalText RetrievalVideo CaptioningVideo Question AnsweringVideo-Text RetrievalVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

2023-09-18 · Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang 외

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to gene…

Image GenerationSurveyText to Image GenerationText-to-Image Generation

Large Multimodal Models: Notes on CVPR 2023 Tutorial

2023-06-26 · Chunyuan Li

This tutorial note summarizes the presentation on ``Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on ``Recent Advances in Vision Foundation Models''. The tutori…

Language ModelingLanguage Modelling

Deep Learning applied to NLP

2017-03-09 · Marc Moreno Lopez, Jugal Kalita

Convolutional Neural Network (CNNs) are typically associated with Computer Vision. CNNs are responsible for major breakthroughs in Image Classification and are the core of most Computer Vision systems today. More recentl…

Deep LearningGeneral Classificationimage-classificationImage Classification

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

2025-07-17 · Hao Sun, Mihaela van der Schaar

In the era of Large Language Models (LLMs), alignment has emerged as a fundamental yet challenging problem in the pursuit of more reliable, controllable, and capable machine intelligence. The recent success of reasoning …

Language ModelingLanguage ModellingLarge Language ModelReinforcement Learning (RL)

Deep-learning PDEs with unlabeled data and hardwiring physics laws

2019-04-13 · S. Mohammad H. Hashemi, Demetri Psaltis

Providing fast and accurate solutions to partial differential equations is a problem of continuous interest to the fields of applied mathematics and physics. With the recent advances in machine learning, the adoption lea…

BIG-bench Machine LearningDeep Learning