Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models
Can vision-language models predict what a 180° rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Prediction: given an original image, a model must answer what would be seen or read after a 180° in-plane rotation, without directly observing the rotated target. To isolate this gap, we introduce RotOutBench, a paired diagnostic benchmark spanning open visual cases and controlled text-image rotations. A sharp pattern emerges: many VLMs can recognize the relevant content when directly given either the original or rotated image, yet fail to infer the rotated result from the original image alone. On controlled text-image rotations, predicted-rotation accuracy collapses to near zero even for models with high direct-reading accuracy. A model-level case study further shows that the prediction state can approach a rotated-image reading state, while the final readout still shifts toward the original string. Current VLMs can recognize a transformed visual state when it is shown, but often fail to predict that state from the original view.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Countering Inconsistent Labelling by Google's Vision API for Rotated Images
Google's Vision API analyses images and provides a variety of output predictions, one such type is context-based labelling. In this paper, it is shown that adversarial examples that cause incorrect label prediction and s…
SPIN: Simplifying Polar Invariance for Neural networks Application to vision-based irradiance forecasting
Translational invariance induced by pooling operations is an inherent property of convolutional neural networks, which facilitates numerous computer vision tasks such as classification. Yet to leverage rotational invaria…
Data AugmentationSolar Irradiance ForecastingOSKDet: Orientation-Sensitive Keypoint Localization for Rotated Object Detection
Rotated object detection is a challenging issue in computer vision field. Inadequate rotated representation and the confusion of parametric regression have been the bottleneck for high performance rotated detection. …
Objectobject-detectionObject DetectionregressionRotation Invariance in Floor Plan Digitization using Zernike Moments
Nowadays, a lot of old floor plans exist in printed form or are stored as scanned raster images. Slight rotations or shifts may occur during scanning. Bringing floor plans of this form into a machine readable form to ena…
FormRAGOSKDet: Towards Orientation-sensitive Keypoint Localization for Rotated Object Detection
Rotated object detection is a challenging issue of computer vision field. Loss of spatial information and confusion of parametric order have been the bottleneck for rotated detection accuracy. In this paper, we propose a…
object-detectionObject Detection