paper-with-me

Papers

MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models

2023-03-23 · CVPR 2023 1 · Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi, Kyoung-Woon On, Byungseok Roh, Hyunwoo J. Kim

Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single task-specific loss is adopted for fine-tuning. However, such fine-tuning methods do not fully leverage other losses that are potentially beneficial for the target task. Therefore, we propose MEta Loss TRansformer (MELTR), a plug-in module that automatically and non-linearly combines various loss functions to aid learning the target task via auxiliary learning. We formulate the auxiliary learning as a bi-level optimization problem and present an efficient optimization algorithm based on Approximate Implicit Differentiation (AID). For evaluation, we apply our framework to various video foundation models (UniVL, Violet and All-in-one), and show significant performance gain on all four downstream tasks: text-to-video retrieval, video question answering, video captioning, and multi-modal sentiment analysis. Our qualitative analyses demonstrate that MELTR adequately transforms' individual loss functions and melts' them into an effective unified loss. Code is available at https://github.com/mlvlab/MELTR.

📄 PDF Abstract BibTeX arXiv:2303.13009

Code (1)

mlvlab/MELTR 공식 구현 pytorch

Tasks

Auxiliary LearningMultimodal Sentiment AnalysisQuestion AnsweringRetrievalSentiment AnalysisText to Video RetrievalTGIF-ActionTGIF-FrameTGIF-TransitionVideo CaptioningVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

CAMELTrack: Context-Aware Multi-cue ExpLoitation for Online Multi-Object Tracking

2025-05-02 · Vladimir Somers, Baptiste Standaert, Victor Joos, Alexandre Alahi 외

Online multi-object tracking has been recently dominated by tracking-by-detection (TbD) methods, where recent advances rely on increasingly sophisticated heuristics for tracklet representation, feature fusion, and multi-…

Multi-Object TrackingObject TrackingOnline Multi-Object Tracking

Empowering Meta-Analysis: Leveraging Large Language Models for Scientific Synthesis

2024-11-16 · Jawad Ibn Ahad, Rafeed Mohammad Sultan, Abraham Kaikobad, Fuad Rahman 외

This study investigates the automation of meta-analysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies support a…

ArticlesPrompt EngineeringRAGRetrieval-augmented Generation

TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting

2023-06-20 · Taorong Liu, Liang Liao, Delin Chen, Jing Xiao 외

Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A re…

DecoderImage InpaintingImage Restoration

Meta-Prior: Meta learning for Adaptive Inverse Problem Solvers

2023-11-30 · Matthieu Terris, Thomas Moreau

Deep neural networks have become a foundational tool for addressing imaging inverse problems. They are typically trained for a specific task, with a supervised loss to learn a mapping from the observations to the image t…

Meta-Learning

Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training

2024-07-17 · Florian Schmid, Paul Primus, Tobias Morocutti, Jonathan Greif 외

This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Challenge. We fine-tune three large Audio …

Event DetectionMissing LabelsSound Event Detection