paper-with-me

Papers

Tutorial on Multimodal Machine Learning

2022-07-01 · NAACL (ACL) 2022 7 · Louis-Philippe Morency, Paul Pu Liang, Amir Zadeh

Multimodal machine learning involves integrating and modeling information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective computing, robotics, finance, HCI, and healthcare. This tutorial, building upon a new edition of a survey paper on multimodal ML as well as previously-given tutorials and academic courses, will describe an updated taxonomy on multimodal machine learning synthesizing its core technical challenges and major directions for future research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningSurvey

Similar Papers 제목 키워드 기반

Multimodal Machine Learning: Integrating Language, Vision and Speech

2017-07-01 · ACL 2017 7 · Louis-Philippe Morency, Tadas Baltru{\v{s}}aitis

Multimodal machine learning is a vibrant multi-disciplinary research field which addresses some of the original goals of artificial intelligence by integrating and modeling multiple communicative modalities, including li…

Audio-Visual Speech RecognitionBIG-bench Machine LearningImage CaptioningQuestion Answering+8

Large Multimodal Models: Notes on CVPR 2023 Tutorial

2023-06-26 · Chunyuan Li

This tutorial note summarizes the presentation on ``Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on ``Recent Advances in Vision Foundation Models''. The tutori…

Language ModelingLanguage Modelling

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

2026-06-02 · Zechen Bai, Zhiheng Chen, Yiqi Lin, Kevin Qinghong Lin 외 arxiv

Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutorial, a framework that transforms this exp…

Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond

2024-10-08 · Soyeon Caren Han, Feiqi Cao, Josiah Poon, Roberto Navigli

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understandi…

Question AnsweringVisual Question AnsweringVisual Storytelling

Recognizing Multimodal Entailment

2021-08-01 · ACL 2021 5 · Cesar Ilharco, Afsaneh Shirazi, Arjun Gopalan, Arsha Nagrani 외

How information is created, shared and consumed has changed rapidly in recent decades, in part thanks to new social platforms and technologies on the web. With ever-larger amounts of unstructured and limited labels, orga…

Graph LearningQuestion Answering