paper-with-me

Papers

MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description

2018-09-19 · Oliver Nina, Washington Garcia, Scott Clouse, Alper Yilmaz

Learning visual feature representations for video analysis is a daunting task that requires a large amount of training samples and a proper generalization framework. Many of the current state of the art methods for video captioning and movie description rely on simple encoding mechanisms through recurrent neural networks to encode temporal visual information extracted from video data. In this paper, we introduce a novel multitask encoder-decoder framework for automatic semantic description and captioning of video sequences. In contrast to current approaches, our method relies on distinct decoders that train a visual encoder in a multitask fashion. Our system does not depend solely on multiple labels and allows for a lack of training data working even with datasets where only one single annotation is viable per video. Our method shows improved performance over current state of the art methods in several metrics on multi-caption and single-caption datasets. To the best of our knowledge, our method is the first method to use a multitask approach for encoding video features. Our method demonstrates its robustness on the Large Scale Movie Description Challenge (LSMDC) 2017 where our method won the movie description task and its results were ranked among other competitors as the most helpful for the visually impaired.

📄 PDF Abstract BibTeX arXiv:1809.07257

Code (1)

OSUPCVLab/VideoToTextDNN 공식 구현 pytorch

Tasks

DecoderVideo Captioning

Similar Papers 제목 키워드 기반

Resting state-fMRI approach towards understanding impairments in mTLE

2020-09-24

Mesial temporal lobe epilepsy (mTLE) is the most common form of epilepsy. While it is characterized by an epileptogenic focus in the mesial temporal lobe, it is increasingly understood as a network disorder. Hence, under…

Deep Asymmetric Multi-task Feature Learning

2017-08-01 · ICML 2018 7 · Hae Beom Lee, Eunho Yang, Sung Ju Hwang

We propose Deep Asymmetric Multitask Feature Learning (Deep-AMTFL) which can learn deep representations shared across multiple tasks while effectively preventing negative transfer that may happen in the feature sharing p…

image-classificationImage ClassificationTransfer Learning

BUSTR: Breast Ultrasound Text Reporting with a Descriptor-Aware Vision-Language Model

2025-11-26 · Rawa Mohammed, Mina Attin, Bryar Shareef arxiv

Automated radiology report generation (RRG) for breast ultrasound (BUS) is limited by the lack of paired image-report datasets and the risk of hallucinations from large language models. We propose BUSTR, a multitask visi…

Multiresolution and Multimodal Speech Recognition with Transformers

2020-04-29 · Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare, Shiva Sundaram

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. W…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Multimodal and Multiresolution Speech Recognition with Transformers

2020-07-01 · ACL 2020 6 · Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare, Shiva Sundaram

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. W…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition