paper-with-me

홈 › Papers

STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision

2025-08-12 · Chen Li, Han Zhang, Zhantao Yang, Fangyi Chen, Zihan Wang, Anudeepsekhar Bolimera, Marios Savvides arxiv

Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM-S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks.

📄 PDF Abstract BibTeX arXiv:2508.08688

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

STELAR: Spatio-temporal Tensor Factorization with Latent Epidemiological Regularization

2020-12-08 · Nikos Kargas, Cheng Qian, Nicholas D. Sidiropoulos, Cao Xiao 외

Accurate prediction of the transmission of epidemic diseases such as COVID-19 is crucial for implementing effective mitigation measures. In this work, we develop a tensor method to predict the evolution of epidemic trend…

AttributePrediction

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

2026-06-25 · Sangjin Han, Hoseong Jung, Jeongtae Her, Changhyun Choi 외 arxiv

Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they of…

Trajectory PredictionAutonomous DrivingMotion Forecasting

LAKe-Net: Topology-Aware Point Cloud Completion by Localizing Aligned Keypoints

2022-03-31 · CVPR 2022 1 · Junshu Tang, Zhijun Gong, Ran Yi, Yuan Xie 외

Point cloud completion aims at completing geometric and topological shapes from a partial observation. However, some topology of the original shape is missing, existing methods directly predict the location of complete p…

Point Cloud Completion

User-Feedback-Driven Adaptation for Vision-and-Language Navigation

2025-12-11 · Yongqiang Yu, Xuhui Li, Hazza Mahmood, Jinxing Zhou 외 arxiv

Real-world deployment of Vision-and-Language Navigation (VLN) agents is constrained by the scarcity of reliable supervision after offline training. While recent adaptation methods attempt to mitigate distribution shifts …

Topology Matters: A Cautionary Case Study of Graph SSL on Neuro-Inspired Benchmarks

2026-02-03 · May Kristine Jonson Carlon, Su Myat Noe, Haojiong Wang, Yasuo Kuniyoshi arxiv

Understanding how local interactions give rise to global brain organization requires models that can represent information across multiple scales. We introduce a hierarchical self-supervised learning (SSL) framework that…

Self-Supervised Learning