Vision Foundation Models · Zhejiang University
VLM vs Video Generation: Which Pretraining Wins at Spatial Intelligence?
A frozen-feature probe pits VLMs against video-generation models on spatial tasks. VLMs win semantics (92.08 mAP vs 69.89), video models win geometry (0.527 camera AUC vs 0.330), and a naive concat of the two beats both.