Papers TGIF-Action
“TGIF-Action” 태그가 달린 논문 7편 · 필터 해제
Lightweight Recurrent Cross-modal Encoder for Video Question Answering
A video question answering task essentially boils down to how to fuse the information between text and video effectively to predict an answer. Most works employ a transformer encoder as a cross-modal encoder to fuse both…
Action RecognitionQuestion AnsweringTGIF-ActionTGIF-Frame+3MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models
Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single ta…
Auxiliary LearningMultimodal Sentiment AnalysisQuestion AnsweringRetrieval+9MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
Video-and-language understanding has a variety of applications in the industry, such as video question answering, text-video retrieval, and multi-label classification. Existing video-and-language understanding methods ge…
Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMultiple-choiceQuestion Answering+7HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training
Video-language pre-training has advanced the performance of various downstream video-language tasks. However, most previous methods directly inherit or adapt typical image-language pre-training paradigms to video-languag…
cross-modal alignmentTGIF-ActionTGIF-FrameTGIF-Transition+6An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling
Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) p…
Fill MaskOptical Flow EstimationQuestion AnsweringRetrieval+8Clover: Towards A Unified Video-Language Alignment and Fusion Model
Building a universal Video-Language model for solving various video understanding tasks (\emph{e.g.}, text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal,…
Language ModelingLanguage ModellingQuestion AnsweringRetrieval+9All in One: Exploring Unified Video-Language Pre-training
Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They pursue better performance via utilizing heav…
AllLanguage ModellingMultiple-choiceQuestion Answering+9