Structuring Space: From Cognition to Intelligence

Project Description

While outdoor spatial perception has advanced rapidly, largely driven by autonomous driving, humans spend most of their time indoors. Indoor environments may be more structured than streets, yet they are in some ways more intricate: layouts are often repetitive, rooms can be hidden or inaccessible, spaces change over time, and reliable GNSS signals are typically unavailable. Geometric 3D reconstruction methods, from point clouds and meshes to NeRFs and 3D Gaussian Splatting, can capture the geometry and appearance of an indoor space, but they do not fully capture what it means, how its different parts are connected, or how it can be used. A 3D scan can show the shape of a room but doesn’t indicate that it is a kitchen, which corridor leads to it, where a person can actually move, or where evacuation routes are accessible.

This project asks: What does an indoor spatial model need beyond geometry to become genuinely useful? Indeed, the representation must capture not only what the environment looks like, but what it is, how it is organized, and how it can be used. Ultimately, this project investigates how to enrich indoor 3D reconstructions with semantic and topological context, moving from raw sensor data toward a unified representation that captures not just geometry, but meaning, structure, and the relationships between places and objects. We envision an integrated spatial intelligence pipeline that enables autonomous systems, such as drones, to reconstruct, reson, refine, localize, and navigate indoor environments independently.

Project Goals

The project follows an indoor scene from raw sensor data to autonomous action, organized around four spatial cognition tasks that build on one another:

  1. Perception: Interpret raw sensor data into a compact, structure-aware understanding of a scene: distinguishing free space from objects, recognizing structure and relationships at a global level, and producing a compact, usable abstraction.

  2. Reconstruction: Translate the extracted spatial information into a single, coherent representation that jointly and consistently encodes geometric structure (what a space looks like), semantic content (what’s in it), and topological organization (how it’s organized).

  3. Localization: Support context-aware, cross-modal localization by leveraging modality-invariant, structure-aware representations, particularly in complex and visually repetitive indoor environments.

  4. Planning and Navigation: Exploit the spatial understanding of the indoor environment to support intelligent, constraint-aware planning and autonomous navigation.

Publications

  1. Mansour, W., & Werner, M. (2024). Calculating Upstream Relation in Spatial Networks Under Path Constraints. Proceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems, 314–324. https://doi.org/10.1145/3678717.3691288 [PDF] [Online]

Acknowledgement

This project is kindly supported by the Technical University of Munich.


Contact

Wejdene Mansour
wejdene.mansour@tum.de
Professorship of Big Geospatial Data Management
Lise-Meitner-Str. 9
85521 Ottobrunn


© 2020 M. Werner | Imprint | Privacy Policy