Context-Dependent Affordance Computation in Vision-Language Models
Abstract
We characterize context-dependent affordance computation in vision-language models. In a study of 3,213 scene-context pairs from COCO-2017 with systematic context priming across 7 agentic personas, Qwen3-VL-30B-A3B exhibits massive affordance drift: mean Jaccard similarity between context conditions is 0.095, indicating that over 90% of lexical scene description is context-dependent; a cross-model replication on LLaVA-1.5-13B reproduces the effect (mean J = 0.160), and sentence-level semantic similarity shows 58.5% context dependence. Stochastic baseline experiments confirm the drift reflects genuine context effects rather than generation noise, and Tucker decomposition reveals stable orthogonal latent factors. We do not claim to establish processing order or architectural primacy; the findings suggest dynamic, query-dependent ontological projection rather than static world modeling as a direction for robotics.