paper-with-me

Papers

How2: A Large-scale Dataset for Multimodal Language Understanding

2018-11-01 · Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, Florian Metze

In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequence-to-sequence baselines for machine translation, automatic speech recognition, spoken language translation, and multimodal summarization. By making available data and code for several multimodal natural language tasks, we hope to stimulate more research on these and similar challenges, to obtain a deeper understanding of multimodality in language processing.

📄 PDF Abstract BibTeX arXiv:1811.00347

Code (2)

srvk/how2-dataset 공식 구현 pytorch
jasonppy/promptingwhisper pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

2024-10-02 · Wenmo Qiu, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded o…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

2023-07-13 · Yi Wang, Yinan He, Yizhuo Li, Kunchang Li 외

This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. The InternVid datase…

Action RecognitionContrastive LearningRepresentation LearningRetrieval+5

MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

2023-11-15 · Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen 외

With the rapid development of large language models (LLMs) and their integration into large multimodal models (LMMs), there has been impressive progress in zero-shot completion of user-oriented vision-language tasks. How…

Chart Understanding

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

2025-03-04 · CVPR 2025 1 · Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram 외

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient …

2D Panoptic SegmentationGraph GenerationLanguage ModelingLanguage Modelling+2

Understanding Chinese Video and Language via Contrastive Multimodal Pre-Training

2021-04-19 · Chenyi Lei, Shixian Luo, Yong liu, Wanggui He 외

The pre-trained neural models have recently achieved impressive performances in understanding multimodal content. However, it is still very challenging to pre-train neural models for video and language understanding, esp…

Contrastive LearningLanguage ModelingLanguage ModellingMasked Language Modeling+1