paper-with-me

Papers

Mounting Video Metadata on Transformer-based Language Model for Open-ended Video Question Answering

2021-08-11 · Donggeon Lee, SeongHo Choi, Youwon Jang, Byoung-Tak Zhang

Video question answering has recently received a lot of attention from multimodal video researchers. Most video question answering datasets are usually in the form of multiple-choice. But, the model for the multiple-choice task does not infer the answer. Rather it compares the answer candidates for picking the correct answer. Furthermore, it makes it difficult to extend to other tasks. In this paper, we challenge the existing multiple-choice video question answering by changing it to open-ended video question answering. To tackle open-ended question answering, we use the pretrained GPT2 model. The model is fine-tuned with video inputs and subtitles. An ablation study is performed by changing the existing DramaQA dataset to an open-ended question answering, and it shows that performance can be improved using video metadata.

📄 PDF Abstract BibTeX arXiv:2108.05158

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMultiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories

2025-02-13 · Tomas Peterka, Matyas Bohacek

The most effective misinformation campaigns are multimodal, often combining text with images and videos taken out of context -- or fabricating them entirely -- to support a given narrative. Contemporary methods for detec…

ArticlesLanguage ModelingLanguage ModellingLarge Language Model+1

Contextual RNN-T For Open Domain ASR

2020-06-04 · Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig 외

End-to-end (E2E) systems for automatic speech recognition (ASR), such as RNN Transducer (RNN-T) and Listen-Attend-Spell (LAS) blend the individual components of a traditional hybrid ASR system - acoustic model, language …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Event and Entity Extraction from Generated Video Captions

2022-11-05 · Johannes Scherer, Ansgar Scherp, Deepayan Bhowmik

Annotation of multimedia data by humans is time-consuming and costly, while reliable automatic generation of semantic metadata is a major challenge. We propose a framework to extract semantic metadata from automatically …

Caption GenerationDense Video CaptioningVideo Captioning

Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model

2020-05-15 · Da-Rong Liu, Chunxi Liu, Frank Zhang, Gabriel Synnaeve 외

Videos uploaded on social media are often accompanied with textual descriptions. In building automatic speech recognition (ASR) systems for videos, we can exploit the contextual information provided by such video metadat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata

2026-02-13 · Prasanna Sridhar, Horace Lee, David M. S. Pinto, Andrew Zisserman 외 arxiv

In this paper, we present WISE, an open-source audiovisual search engine which integrates a range of multimodal retrieval capabilities into a single, practical tool accessible to users without machine learning expertise.…