paper-with-me

홈 › Papers

Omnivore: A Single Model for Many Visual Modalities

2022-01-20 · CVPR 2022 1 · Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, Ishan Misra

Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data. Instead, in this paper, we propose a single model which excels at classifying images, videos, and single-view 3D data using exactly the same model parameters. Our 'Omnivore' model leverages the flexibility of transformer-based architectures and is trained jointly on classification tasks from different modalities. Omnivore is simple to train, uses off-the-shelf standard datasets, and performs at-par or better than modality-specific models of the same size. A single Omnivore model obtains 86.0% on ImageNet, 84.1% on Kinetics, and 67.1% on SUN RGB-D. After finetuning, our models outperform prior work on a variety of vision tasks and generalize across modalities. Omnivore's shared visual representation naturally enables cross-modal recognition without access to correspondences between modalities. We hope our results motivate researchers to model visual modalities together.

📄 PDF Abstract BibTeX arXiv:2201.08377

Code (2)

facebookresearch/omnivore 공식 구현 pytorch
towhee-io/towhee pytorch

Tasks

Action ClassificationAction RecognitionImage ClassificationmodelScene RecognitionSemantic Segmentation

Similar Papers 제목 키워드 기반

Emu: Generative Pretraining in Multimodality

2023-07-11 · Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang 외

We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscri…

Image CaptioningImage GenerationImage to textQuestion Answering+7

A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge

2022-11-16 · Sicheng Mo, Fangzhou Mu, Yin Li

This report describes Badgers@UW-Madison, our submission to the Ego4D Natural Language Queries (NLQ) Challenge. Our solution inherits the point-based event representation from our prior work on temporal action localizati…

Action LocalizationNatural Language QueriesTemporal Action LocalizationVideo Grounding

Flexible-modal Deception Detection with Audio-Visual Adapter

2023-02-11 · Zhaoxu Li, Zitong Yu, Nithish Muthuchamy Selvaraj, Xiaobao Guo 외

Detecting deception by human behaviors is vital in many fields such as custom security and multimedia anti-fraud. Recently, audio-visual deception detection attracts more attention due to its better performance than usin…

Deception Detection

Deep Learning for Multi-Task Medical Image Segmentation in Multiple Modalities

2017-04-11 · Pim Moeskops, Jelmer M. Wolterink, Bas H. M. van der Velden, Kenneth G. A. Gilhuijs 외

Automatic segmentation of medical images is an important task for many clinical applications. In practice, a wide range of anatomical structures are visualised using different imaging modalities. In this paper, we invest…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Where a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries Challenge

2022-11-16 · Fangzhou Mu, Sicheng Mo, Gillian Wang, Yin Li

This report describes our submission to the Ego4D Moment Queries Challenge 2022. Our submission builds on ActionFormer, the state-of-the-art backbone for temporal action localization, and a trio of strong video features …

Action LocalizationMoment QueriesTemporal Action Localization