Temporal/Casual QA
1개 벤치마크 · 논문 6편 · 이 태스크의 논문 보기 →
Benchmarks
NExT-QA
Most implemented
Flamingo: a Visual Language Model for Few-Shot Learning
Emu: Generative Pretraining in Multimodality
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Gemini: A Family of Highly Capable Multimodal Models
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Papers
Gemini: A Family of Highly Capable Multimodal Models
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitabl…
1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationImage Retrieval+5PaLI-3 Vision Language Models: Smaller, Faster, Stronger
This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this strong performance, we compare Vision Tra…
Chart Question AnsweringCross-Modal Retrievalimage-classificationImage Classification+6Emu: Generative Pretraining in Multimodality
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscri…
Image CaptioningImage GenerationImage to textQuestion Answering+7Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models
Video Question Answering (VideoQA) has been significantly advanced from the scaling of recent Large Language Models (LLMs). The key idea is to convert the visual information into the language feature space so that the ca…
cross-modal alignmentDomain GeneralizationQuestion AnsweringRetrieval+2PaLI-X: On Scaling up a Multilingual Vision and Language Model
We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new leve…
Chart Question Answeringdocument understandingFine-Grained Image RecognitionIn-Context Learning+10Flamingo: a Visual Language Model for Few-Shot Learning
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Mode…
Few-Shot LearningGenerative Visual Question AnsweringLanguage ModelingLanguage Modelling+12