paper-with-me

홈 › Papers

OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models

2022-12-08 · Jinze Bai, Rui Men, Hao Yang, Xuancheng Ren, Kai Dang, Yichang Zhang, Xiaohuan Zhou, Peng Wang, Sinan Tan, An Yang, Zeyu Cui, Yu Han, Shuai Bai, Wenbin Ge, Jianxin Ma, Junyang Lin, Jingren Zhou, Chang Zhou

Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alternative to approaching general-purpose AI, existing generalist models are still at an early stage, where modality and task coverage is limited. To empower multi-modal task-scaling and speed up this line of research, we release a generalist model learning system, OFASys, built on top of a declarative task interface named multi-modal instruction. At the core of OFASys is the idea of decoupling multi-modal task representations from the underlying model implementations. In OFASys, a task involving multiple modalities can be defined declaratively even with just a single line of code. The system automatically generates task plans from such instructions for training and inference. It also facilitates multi-task training for diverse multi-modal workloads. As a starting point, we provide presets of 7 different modalities and 23 highly-diverse example tasks in OFASys, with which we also develop a first-in-kind, single model, OFA+, that can handle text, image, speech, video, and motion data. The single OFA+ model achieves 95% performance in average with only 16% parameters of 15 task-finetuned models, showcasing the performance reliability of multi-modal task-scaling provided by OFASys. Available at https://github.com/OFA-Sys/OFASys

📄 PDF Abstract BibTeX arXiv:2212.04408

Code (1)

ofa-sys/ofasys 공식 구현 pytorch

Tasks

Multi-Task Learning

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models

2024-08-26 · Ian Stewart, Sameera Horawalavithana, Brendan Kennedy, Sai Munikoti 외

Multimodal foundation models (MFMs) such as OFASys show the potential to unlock analysis of complex data such as images, videos, and audio data via text prompts alone. However, their performance may suffer in the face of…

Data Augmentation

Learning to Embed Multi-Modal Contexts for Situated Conversational Agents

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of…

coreference-resolutionCoreference ResolutionDecoderdialog state tracking+2

Learning to Embed Multi-Modal Contexts for Situated Conversational Agents

2022-07-01 · Findings (NAACL) 2022 7 · Haeju Lee, Oh Joon Kwon, Yunseon Choi, Minho Park 외

The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of…

coreference-resolutionCoreference ResolutionDecoderdialog state tracking+3

Mutual Information Analysis in Multimodal Learning Systems

2024-05-21 · Hadi Hadizadeh, S. Faegheh Yeganli, Bahador Rashidi, Ivan V. Bajić

In recent years, there has been a significant increase in applications of multimodal signal processing and analysis, largely driven by the increased availability of multimodal datasets and the rapid progress in multimoda…

3D Object DetectionAutonomous DrivingAutonomous Vehiclesobject-detection+1

Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description

2017-10-19 · WS 2017 9 · Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares 외

We present the results from the second shared task on multimodal machine translation and multilingual image description. Nine teams submitted 19 systems to two tasks. The multimodal translation task, in which the source …

Image DescriptionMachine TranslationMultimodal Machine TranslationSentence+1