paper-with-me

Papers

AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation

2025-08-04 · Ziyang Luo, Nian Liu, Fahad Shahbaz Khan, Junwei Han arxiv

Reference Audio-Visual Segmentation (Ref-AVS) tasks challenge models to precisely locate sounding objects by integrating visual, auditory, and textual cues. Existing methods often lack genuine semantic understanding, tending to memorize fixed reasoning patterns. Furthermore, jointly training for reasoning and segmentation can compromise pixel-level precision. To address these issues, we introduce AURORA, a novel framework designed to enhance genuine reasoning and language comprehension in reference audio-visual segmentation. We employ a structured Chain-of-Thought (CoT) prompting mechanism to guide the model through a step-by-step reasoning process and introduce a novel segmentation feature distillation loss to effectively integrate these reasoning abilities without sacrificing segmentation performance. To further cultivate the model's genuine reasoning capabilities, we devise a further two-stage training strategy: first, a ``corrective reflective-style training" stage utilizes self-correction to enhance the quality of reasoning paths, followed by reinforcement learning via Group Reward Policy Optimization (GRPO) to bolster robustness in challenging scenarios. Experiments demonstrate that AURORA achieves state-of-the-art performance on Ref-AVS benchmarks and generalizes effectively to unreferenced segmentation.

📄 PDF Abstract BibTeX arXiv:2508.02149

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Aurora: Neuro-Symbolic AI Driven Advising Agent

2026-02-20 · Lorena Amanda Quincoso Lugones, Christopher Kverne, Nityam Sharadkumar Bhimani, Ana Carolina Oliveira 외 arxiv

Academic advising in higher education is under severe strain, with advisor-to-student ratios commonly exceeding 300:1. These structural bottlenecks limit timely access to guidance, increase the risk of delayed graduation…

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

2024-12-04 · CVPR 2025 1 · Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen 외

Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning ab…

Depth Estimationobject-detectionObject DetectionVisual Reasoning

Aurora: Unified Video Editing with a Tool-Using Agent

2026-05-18 · Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao, Zhenghong Zhou 외 arxiv

Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set of weights covers replacement, removal, s…

Style Transfer

Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

2026-08-21 · Yitao Xu, Tong Wu, Yiyan Wu, Guoji Xu 외 arxiv

Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and la…

Trajectory PlanningAutonomous Driving

Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

2024-07-03 · Benno Krojer, Dheeraj Vattikonda, Luis Lara, Varun Jampani 외

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general ins…

AttributeSpatial Reasoning