paper-with-me

Papers

M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis

2023-10-23 · Fei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang, Jianbing Zhang, Xinyu Dai

Multimodal Aspect-based Sentiment Analysis (MABSA) is a fine-grained Sentiment Analysis task, which has attracted growing research interests recently. Existing work mainly utilizes image information to improve the performance of MABSA task. However, most of the studies overestimate the importance of images since there are many noise images unrelated to the text in the dataset, which will have a negative impact on model learning. Although some work attempts to filter low-quality noise images by setting thresholds, relying on thresholds will inevitably filter out a lot of useful image information. Therefore, in this work, we focus on whether the negative impact of noisy images can be reduced without modifying the data. To achieve this goal, we borrow the idea of Curriculum Learning and propose a Multi-grained Multi-curriculum Denoising Framework (M2DF), which can achieve denoising by adjusting the order of training data. Extensive experimental results show that our framework consistently outperforms state-of-the-art work on three sub-tasks of MABSA.

📄 PDF Abstract BibTeX arXiv:2310.14605

Code (1)

grandchicken/m2df 공식 구현 pytorch

Tasks

Aspect-Based Sentiment AnalysisDenoisingSentiment Analysis

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A Dual-Module Denoising Approach with Curriculum Learning for Enhancing Multimodal Aspect-Based Sentiment Analysis

2024-12-11 · Nguyen Van Doan, Dat Tran Nguyen, Cam-Van Thi Nguyen

Multimodal Aspect-Based Sentiment Analysis (MABSA) combines text and images to perform sentiment analysis but often struggles with irrelevant or misleading visual information. Existing methodologies typically address eit…

Aspect-Based Sentiment AnalysisDenoisingImage DenoisingSentence+1

Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

2025-09-29 · Yuanshuai Li, Yuping Yan, Junfeng Tang, Yunxuan Li 외 arxiv

Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evide…

DreamReasoner-8B: Block-Size Curriculum Learning for Diffusion Reasoning Models

2026-06-17 · Zirui Wu, Lin Zheng, Jiacheng Ye, Shansan Gong 외 arxiv

Block diffusion language models accelerate decoding through parallel block-wise denoising, yet whether they can be reliably scaled for long chain-of-thought (CoT) reasoning remains unresolved. To this end, we develop Dre…

Multi-Task Curriculum Transfer Deep Learning of Clothing Attributes

2016-10-12 · Qi Dong, Shaogang Gong, Xiatian Zhu

Recognising detailed clothing characteristics (fine-grained attributes) in unconstrained images of people in-the-wild is a challenging task for computer vision, especially when there is only limited training data from th…

AttributeDeep LearningTransfer Learning

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

2025-06-01 · Zhengcong Fei, Hao Jiang, Di Qiu, Baoxuan Gu 외

The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present SkyReels-Audio, a unified framework for…

Denoising