Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception
Abstract
Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs reproduce several human-like biases,whilesupervisedCNNsdonot.DoVision-LanguageModels(VLMs) exhibit similar competences? Across multiple VLM families and model scales,zero-shotandin-contextpromptingbothproducedistinctivefailures: slantispredictedatonlyasmallsetofanchors(e.g.,0°,±25°,±45°)withlittle dependence on stimulus field of view, optical slant, or surface curvature. Supervised fine-tuning partially remediates the failure, but residual anchoring persists. While success in high-level vision-language benchmarks might notrequiresensitivitytolow-levelgeometriccues,weinterpretanchoringas a failure at the representation-to-output language interface: not necessarily anabsenceofgeometricencoding,butafailuretoexpressitinagradedform.