Fine-tuning and Sampling Strategies for Multimodal Role Labeling of Entities under Class Imbalance
We propose our solution to the multimodal semantic role labeling task from the CONSTRAINT’22 workshop. The task aims at classifying entities in memes into classes such as “hero” and “villain”. We use several pre-trained multi-modal models to jointly encode the text and image of the memes, and implement three systems to classify the role of the entities. We propose dynamic sampling strategies to tackle the issue of class imbalance. Finally, we perform qualitative analysis on the representations of the entities.
Code (0)
등록된 구현이 없습니다.
Tasks
Semantic Role LabelingSimilar Papers 제목 키워드 기반
Latent Space Guided Scenario Sampling for Multimodal Segmentation Under Missing Modalities
Multimodal semantic segmentation benefits remote sensing analysis by combining complementary information from different sensor modalities. In real-world remote sensing applications, one or more modalities may be unavaila…
Semantic SegmentationIterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks. Existing approaches for training these a…
Large Language ModelA Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets
We introduce a scaling law for fine-tuning large language models (LLMs) under fixed compute budgets that explicitly accounts for data composition. Conventional approaches measure training data solely by total tokens, yet…
MMLUVideo2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic…
Beyond Fine-Tuning: A Systematic Study of Sampling Techniques in Personalized Image Generation
Personalized text-to-image generation aims to create images tailored to user-defined concepts and textual descriptions. Balancing the fidelity of the learned concept with its ability for generation in various contexts pr…
Image GenerationPersonalized Image GenerationText to Image GenerationText-to-Image Generation