Artificial Intelligence #artificial intelligence#multimodal
New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs
Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.
Jul 8, 2026 1 source