11am - 12 noon
Wednesday 14 October 2026
Toward Spatial Intelligence with Foundation Models: Perception, Generation, and Reconstruction
PhD Viva Open Presentation - Haosen Yang
Hybrid Meeting (21BA02 & Teams) - All Welcome!
Free
University of Surrey
Guildford
Surrey
GU2 7XH
Toward Spatial Intelligence with Foundation Models: Perception, Generation, and Reconstruction
Abstract:
Spatial intelligence enables visual systems to understand how objects, regions, and scenes are organised in the physical world. It spans multiple levels of visual understanding: identifying meaningful regions and their semantics, preserving coherent structures during generation, inferring geometry from individual observations, and relating multiple views to a consistent 3D scene. Recent foundation models provide powerful semantic and generative priors learned from large-scale data, yet these priors are not inherently structured to capture such spatial information. This thesis investigates how foundation-model priors can be adapted to support spatially grounded visual understanding across perception, generation, and reconstruction. To this end, it presents four contributions that progressively advance from region-level understanding to image generation, monocular geometry, and multi-view 3D reconstruction.
The first contribution addresses spatial intelligence in perception through open-world region-level understanding. Recognize Any Region combines the complementary capabilities of SAM and CLIP to localise meaningful regions and recognise their semantic content. By aligning position-aware regional features with pretrained semantic knowledge, it enables flexible region recognition without large-scale task-specific training.
The second contribution studies spatial intelligence in high-resolution image generation. FAM Diffusion identifies mismatches in frequency and attention dynamics when pretrained diffusion models generate images beyond their native resolution. It introduces inference-time frequency and attention modulation to preserve global layouts and local textures across scales, enabling coherent high-resolution generation without additional training.
The third contribution extends spatial understanding from visual appearance to monocular geometry. GeoNeXt repurposes pretrained video diffusion models for joint depth and surface-normal estimation by formulating geometry prediction as a generative next-frame prediction problem. This formulation exploits the structural and temporal priors encoded by video generation models, demonstrating their potential to recover the geometric structure underlying a single image.
The final contribution advances from monocular geometry to explicit 3D scene reconstruction. Localized Points Management improves Gaussian Splatting by regulating the distribution and organisation of its 3D primitives. It combines rendering errors, geometric matching, and multi-view constraints to identify under-reconstructed regions and ill-conditioned primitives, improving geometric fidelity and optimisation stability.
Together, these contributions establish a progressive framework for spatial intelligence, moving from semantic regions and coherent image structures to single-view geometry and multi-view 3D scenes. Across these stages, foundation-model priors are aligned, modulated, repurposed, and structurally constrained according to the spatial requirements of each problem. The thesis thereby advances visual systems that are more data-efficient, controllable, geometry-aware, and spatially grounded.
