paper-with-me

홈 › Papers

VPAI_Lab at MedVidQA 2022: A Two-Stage Cross-modal Fusion Method for Medical Instructional Video Classification

2022-05-01 · BioNLP (ACL) 2022 5 · Bin Li, Yixuan Weng, Fei Xia, Bin Sun, Shutao Li

This paper introduces the approach of VPAI_Lab team’s experiments on BioNLP 2022 shared task 1 Medical Video Classification (MedVidCL). Given an input video, the MedVidCL task aims to correctly classify it into one of three following categories: Medical Instructional, Medical Non-instructional, and Non-medical. Inspired by its dataset construction process, we divide the classification process into two stages. The first stage is to classify videos into medical videos and non-medical videos. In the second stage, for those samples classified as medical videos, we further classify them into instructional videos and non-instructional videos. In addition, we also propose the cross-modal fusion method to solve the video classification, such as fusing the text features (question and subtitles) from the pre-training language models and visual features from image frames. Specifically, we use textual information to concatenate and query the visual information for obtaining better feature representation. Extensive experiments show that the proposed method significantly outperforms the official baseline method by 15.4% in the F1 score, which shows its effectiveness. Finally, the online results show that our method ranks the Top-1 on the online unseen test set. All the experimental codes are open-sourced at https://github.com/Lireanstar/MedVidCL.

📄 PDF Abstract BibTeX

Code (1)

lireanstar/medvidcl 공식 구현 pytorch

Tasks

Video Classification

Similar Papers 제목 키워드 기반

AdvPaint: Protecting Images from Inpainting Manipulation via Adversarial Attention Disruption

2025-03-13 · Joonsung Jeon, Woo Jae Kim, Suhyeon Ha, Sooel Son 외

The outstanding capability of diffusion models in generating high-quality images poses significant threats when misused by adversaries. In particular, we assume malicious adversaries exploiting diffusion models for inpai…

Image Generation

MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D

2024-11-04 · CVPR 2025 1 · Wei Cheng, Juncheng Mu, Xianfang Zeng, Xin Chen 외

Texturing is a crucial step in the 3D asset production workflow, which enhances the visual appeal and diversity of 3D assets. Despite recent advancements in Text-to-Texture (T2T) generation, existing methods often yield …

3D InpaintingSuper-Resolution

MVPainter: Accurate and Detailed 3D Texture Generation via Multi-View Diffusion with Geometric Control

2025-05-19 · Mingqi Shao, Feng Xiong, Zhaoxu Sun, Mu Xu

Recently, significant advances have been made in 3D object generation. Building upon the generated geometry, current pipelines typically employ image diffusion models to generate multi-view RGB images, followed by UV tex…

3D geometryTexture Synthesis

A Dataset for Medical Instructional Video Classification and Question Answering

2022-01-30 · Deepak Gupta, Kush Attal, Dina Demner-Fushman

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may pr…

ClassificationQuestion AnsweringVideo ClassificationVideo Understanding

Overview of the MedVidQA 2022 Shared Task on Medical Video Question-Answering

2022-05-01 · BioNLP (ACL) 2022 5 · Deepak Gupta, Dina Demner-Fushman

In this paper, we present an overview of the MedVidQA 2022 shared task, collocated with the 21st BioNLP workshop at ACL 2022. The shared task addressed two of the challenges faced by medical video question answering: (I)…

Question AnsweringVideo ClassificationVideo Question AnsweringVideo Understanding