paper-with-me

Papers

MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound

2022-01-07 · CVPR 2022 1 · Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, Yejin Choi

As humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce MERLOT Reserve, a model that represents videos jointly over time -- through a new training objective that learns from audio, subtitles, and video frames. Given a video, we replace snippets of text and audio with a MASK token; the model learns by choosing the correct masked-out snippet. Our objective learns faster than alternatives, and performs well at scale: we pretrain on 20 million YouTube videos. Empirical results show that MERLOT Reserve learns strong multimodal representations. When finetuned, it sets state-of-the-art on Visual Commonsense Reasoning (VCR), TVQA, and Kinetics-600; outperforming prior work by 5%, 7%, and 1.5% respectively. Ablations show that these tasks benefit from audio pretraining -- even VCR, a QA task centered around images (without sound). Moreover, our objective enables out-of-the-box prediction, revealing strong multimodal commonsense understanding. In a fully zero-shot setting, our model obtains competitive results on four video tasks, even outperforming supervised approaches on the recently proposed Situated Reasoning (STAR) benchmark. We analyze why audio enables better vision-language representations, suggesting significant opportunities for future research. We conclude by discussing ethical and societal implications of multimodal pretraining.

📄 PDF Abstract BibTeX arXiv:2201.02639

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationNavigateVideo UnderstandingVisual Commonsense Reasoning

Similar Papers 제목 키워드 기반

MERLOT: Multimodal Neural Script Knowledge Models

2021-06-04 · NeurIPS 2021 12 · Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 외

As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce MERLOT, a model that learns multimodal sc…

Multimodal ReasoningVisual Commonsense Reasoning

MERLOT: A Distilled LLM-based Mixture-of-Experts Framework for Scalable Encrypted Traffic Classification

2024-11-20 · Yuxuan Chen, Rongpeng Li, Zhifeng Zhao, Honggang Zhang

We present MERLOT, a scalable mixture-of-expert (MoE) based refinement of distilled large language model optimized for encrypted traffic classification. By applying model distillation techniques in a teacher-student para…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2

Temporal information extraction from clinical text

2017-04-01 · EACL 2017 4 · Julien Tourille, Olivier Ferret, Xavier Tannier, Aur{\'e}lie N{\'e}v{\'e}ol

In this paper, we present a method for temporal relation extraction from clinical narratives in French and in English. We experiment on two comparable corpora, the MERLOT corpus and the THYME corpus, and show that a comm…

RelationRelation ExtractionTemporal Information ExtractionTemporal Relation Extraction

Video Pre-trained Transformer: A Multimodal Mixture of Pre-trained Experts

2023-03-24 · Kastan Day, Daniel Christl, Rohan Salvi, Pranav Sriram

We present Video Pre-trained Transformer. VPT uses four SOTA encoder models from prior work to convert a video into a sequence of compact embeddings. Our backbone, based on a reference Flan-T5-11B architecture, learns a …

Causal Language ModelingLanguage ModelingLanguage Modelling

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

2026-03-30 · Luigi Curini, Alfio Ferrara, Giovanni Pagano, Sergio Picascia arxiv

Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to transcribe Italian parliamentary speec…

Speaker IdentificationSemantic SegmentationEntity Linking