paper-with-me

Papers

SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning

2026-03-05 · Zhu Li, Yongjian Chen, Huiyuan Lai, Xiyuan Gao, Shekhar Nayak, Matt Coler arxiv

Multimodal sarcasm detection requires resolving pragmatic incongruity across textual, acoustic, and visual cues through cross-modal reasoning. To enable robust sarcasm reasoning with foundation models, we propose SarcasmMiner, a reinforcement learning based post-training framework that resists hallucination in multimodal reasoning. We reformulate sarcasm detection as structured reasoning and adopt a dual-track distillation strategy: high-quality teacher trajectories initialize the student model, while the full set of trajectories trains a generative reward model (GenRM) to evaluate reasoning quality. The student is optimized with group relative policy optimization (GRPO) using decoupled rewards for accuracy and reasoning quality. On MUStARD++, SarcasmMiner increases F1 from 59.83% (zero-shot), 68.23% (supervised finetuning) to 70.22%. These findings suggest that reasoning-aware reward modeling enhances both performance and multimodal grounding.

📄 PDF Abstract BibTeX arXiv:2603.05275

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMultimodal ReasoningSarcasm Detection

Similar Papers 제목 키워드 기반

3D-MuPPET: 3D Multi-Pigeon Pose Estimation and Tracking

2023-08-29 · Urs Waldmann, Alex Hoi Hang Chan, Hemal Naik, Máté Nagy 외

Markerless methods for animal posture tracking have been rapidly developing recently, but frameworks and benchmarks for tracking large animal groups in 3D are still lacking. To overcome this gap in the literature, we pre…

Pose Estimation

Tracking the Spatiotemporal Evolution of Landslide Scars Using a Vision Foundation Model: A Novel and Universal Framework

2025-10-11 · Meijun Zhou, Gang Mei, Zhengjing Ma, Nengxiong Xu 외 arxiv

Tracking the spatiotemporal evolution of large-scale landslide scars is critical for understanding the evolution mechanisms and failure precursors, enabling effective early-warning. However, most existing studies have fo…

Video Segmentation

Post-Completion Learning for Language Models

2025-07-27 · Xiang Fei, Siqi Wang, Shu Wei, Yuxiang Nie 외 arxiv

Current language model training paradigms typically terminate learning upon reaching the end-of-sequence (<eos>) token, overlooking the potential learning opportunities in the post-completion space. We propose Post-Compl…

Reinforcement Learning

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

2026-06-29 · Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu 외 arxiv

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mix…

TransFiner: A Full-Scale Refinement Approach for Multiple Object Tracking

2022-07-26 · Bin Sun

Multiple object tracking (MOT) is the task containing detection and association. Plenty of trackers have achieved competitive performance. Unfortunately, for the lack of informative exchange on these subtasks, they are o…

DecoderMultiple Object TrackingObject Tracking