paper-with-me

Papers

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

2025-12-02 · Weiqi Li, Quande Zhang, Ruifeng Zhai, Liang Lin, Guangrun Wang arxiv

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose a one-shot adaptation framework that recalibrates visual representations through lightweight, learnable updates. Our first method, Feature Token Modulation (FTM), applies a global affine transformation to visual tokens and improves Libero viewpoint accuracy from 48.5% to 87.1% with only 4K parameters. Building on this, Feature Linear Adaptation (FLA) introduces low-rank updates to the ViT encoder, achieving 90.8% success with 4.7M parameters -- matching LoRA-scale finetuning at far lower cost. Together, these results reveal substantial untapped robustness in pretrained VLA models and demonstrate that targeted, minimal visual adaptation is sufficient to restore viewpoint generalization.

📄 PDF Abstract BibTeX arXiv:2512.02902

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RESCHED: Rethinking Flexible Job Shop Scheduling from a Transformer-based Architecture with Simplified States

2026-03-07 · Xiangjie Xiao, Cong Zhang, Wen Song, Zhiguang Cao arxiv

Neural approaches to the Flexible Job Shop Scheduling Problem (FJSP), particularly those based on deep reinforcement learning (DRL), have gained growing attention in recent years. However, existing methods rely on comple…

Reinforcement Learning

Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization

2025-11-27 · Yifan Du, Kun Zhou, Yingqian Min, Yue Ling 외 arxiv

We study how different Chain-of-Thought (CoT) designs affect the acquisition of the generalizable visual reasoning ability in vision-language models (VLMs). While CoT data, especially long or visual CoT such as "think wi…

Visual Reasoning

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking

2025-05-25 · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

Recent advances in large language models (LLMs) demonstrate their impressive reasoning capabilities. However, the reasoning confined to internal parametric space limits LLMs' access to real-time information and understan…

Mathematical ReasoningMulti-hop Question AnsweringQuestion Answeringtext-based games

Revisiting Physical-World Adversarial Attack on Traffic Sign Recognition: A Commercial Systems Perspective

2024-09-15 · Ningfei Wang, Shaoyuan Xie, Takami Sato, Yunpeng Luo 외

Traffic Sign Recognition (TSR) is crucial for safe and correct driving automation. Recent works revealed a general vulnerability of TSR models to physical-world adversarial attacks, which can be low-cost, highly deployab…

Adversarial AttackMemorizationTraffic Sign Recognition

Rethink, Revisit, Revise: A Spiral Reinforced Self-Revised Network for Zero-Shot Learning

2021-12-01 · Zhe Liu, Yun Li, Lina Yao, Julian McAuley 외

Current approaches to Zero-Shot Learning (ZSL) struggle to learn generalizable semantic knowledge capable of capturing complex correlations. Inspired by \emph{Spiral Curriculum}, which enhances learning processes by revi…

AttributeZero-Shot Learning