Omnivore: A Single Model for Many Visual Modalities
Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data. Instead, in this paper, we propose a single model which excels at classifying images, videos, and single-view 3D data using exactly the same model parameters. Our 'Omnivore' model leverages the flexibility of transformer-based architectures and is trained jointly on classification tasks from different modalities. Omnivore is simple to train, uses off-the-shelf standard datasets, and performs at-par or better than modality-specific models of the same size. A single Omnivore model obtains 86.0% on ImageNet, 84.1% on Kinetics, and 67.1% on SUN RGB-D. After finetuning, our models outperform prior work on a variety of vision tasks and generalize across modalities. Omnivore's shared visual representation naturally enables cross-modal recognition without access to correspondences between modalities. We hope our results motivate researchers to model visual modalities together.
Code (2)
Tasks
Action ClassificationAction RecognitionImage ClassificationmodelScene RecognitionSemantic SegmentationSimilar Papers 제목 키워드 기반
Emu: Generative Pretraining in Multimodality
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscri…
Image CaptioningImage GenerationImage to textQuestion Answering+7A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge
This report describes Badgers@UW-Madison, our submission to the Ego4D Natural Language Queries (NLQ) Challenge. Our solution inherits the point-based event representation from our prior work on temporal action localizati…
Action LocalizationNatural Language QueriesTemporal Action LocalizationVideo GroundingFlexible-modal Deception Detection with Audio-Visual Adapter
Detecting deception by human behaviors is vital in many fields such as custom security and multimedia anti-fraud. Recently, audio-visual deception detection attracts more attention due to its better performance than usin…
Deception DetectionDeep Learning for Multi-Task Medical Image Segmentation in Multiple Modalities
Automatic segmentation of medical images is an important task for many clinical applications. In practice, a wide range of anatomical structures are visualised using different imaging modalities. In this paper, we invest…
Image SegmentationMedical Image SegmentationSegmentationSemantic SegmentationWhere a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries Challenge
This report describes our submission to the Ego4D Moment Queries Challenge 2022. Our submission builds on ActionFormer, the state-of-the-art backbone for temporal action localization, and a trio of strong video features …
Action LocalizationMoment QueriesTemporal Action Localization