paper-with-me

Papers TGIF-Transition

“TGIF-Transition” 태그가 달린 논문 7편 · 필터 해제

Lightweight Recurrent Cross-modal Encoder for Video Question Answering

2023-06-30 · Knowledge-Based Systems 2023 6 · Steve Andreas Immanuel, Cheol Jeong

A video question answering task essentially boils down to how to fuse the information between text and video effectively to predict an answer. Most works employ a transformer encoder as a cross-modal encoder to fuse both…

Action RecognitionQuestion AnsweringTGIF-ActionTGIF-Frame+3

MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models

2023-03-23 · CVPR 2023 1 · Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi, Kyoung-Woon On 외

Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single ta…

Auxiliary LearningMultimodal Sentiment AnalysisQuestion AnsweringRetrieval+9

MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling

2023-03-10 · Jiaqi Xu, Bo Liu, Yunkuo Chen, Mengli Cheng 외

Video-and-language understanding has a variety of applications in the industry, such as video question answering, text-video retrieval, and multi-label classification. Existing video-and-language understanding methods ge…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMultiple-choiceQuestion Answering+7

HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

2022-12-30 · ICCV 2023 1 · Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu 외

Video-language pre-training has advanced the performance of various downstream video-language tasks. However, most previous methods directly inherit or adapt typical image-language pre-training paradigms to video-languag…

cross-modal alignmentTGIF-ActionTGIF-FrameTGIF-Transition+6

An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

2022-09-04 · CVPR 2023 1 · Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin 외

Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) p…

Fill MaskOptical Flow EstimationQuestion AnsweringRetrieval+8

Clover: Towards A Unified Video-Language Alignment and Fusion Model

2022-07-16 · CVPR 2023 1 · Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu 외

Building a universal Video-Language model for solving various video understanding tasks (\emph{e.g.}, text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal,…

Language ModelingLanguage ModellingQuestion AnsweringRetrieval+9

All in One: Exploring Unified Video-Language Pre-training

2022-03-14 · CVPR 2023 1 · Alex Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge 외

Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They pursue better performance via utilizing heav…

AllLanguage ModellingMultiple-choiceQuestion Answering+9
1–7 / 7