paper-with-me

Papers Temporal/Casual QA

“Temporal/Casual QA” 태그가 달린 논문 6편 · 필터 해제

Gemini: A Family of Highly Capable Multimodal Models

2023-12-19 · The Keyword 2023 12 · Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac 외

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitabl…

1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationImage Retrieval+5

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

2023-10-13 · Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov 외

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this strong performance, we compare Vision Tra…

Chart Question AnsweringCross-Modal Retrievalimage-classificationImage Classification+6

Emu: Generative Pretraining in Multimodality

2023-07-11 · Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang 외

We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscri…

Image CaptioningImage GenerationImage to textQuestion Answering+7

Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models

2023-06-15 · Junting Pan, Ziyi Lin, Yuying Ge, Xiatian Zhu 외

Video Question Answering (VideoQA) has been significantly advanced from the scaling of recent Large Language Models (LLMs). The key idea is to convert the visual information into the language feature space so that the ca…

cross-modal alignmentDomain GeneralizationQuestion AnsweringRetrieval+2

PaLI-X: On Scaling up a Multilingual Vision and Language Model

2023-05-29 · Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa 외

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new leve…

Chart Question Answeringdocument understandingFine-Grained Image RecognitionIn-Context Learning+10

Flamingo: a Visual Language Model for Few-Shot Learning

2022-04-29 · DeepMind 2022 4 · Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 외

Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Mode…

Few-Shot LearningGenerative Visual Question AnsweringLanguage ModelingLanguage Modelling+12
1–6 / 6