An AI model can identify pixels that appear to belong to a roof or building. That result is useful, but it is not yet a building that a project team can track, query or update.

Semantic point cloud of a building and its surroundings, with classified points retained as evidence

*Classified points are evidence. A building entity requires persistent identity, relationships and review.*

Table of contents

The temptation of the mask

Segmentation output looks decisive. A colored polygon appears over an aerial image, its area can be calculated, and it can be exported to GIS. It is tempting to call that polygon “the building.” Yet the mask only describes what the model saw in one source image under one set of conditions.

Shadows, trees, scaffolding, temporary structures, roof equipment and occlusion can change the result. A second flight may produce a different outline even when the physical building did not change. If every detection becomes a new object, the database fills with duplicates. If detections overwrite one another, project history disappears.

<figure><img loading="lazy" decoding="async" width="1400" height="893" src="/media/insights/semantic-mask-multiview.webp" alt="Drone image of roofs beside proposed semantic masks for buildings, vegetation and equipment"><figcaption>Skylens experiment output: semantic masks are proposed observations from one image. They are not verified building identities and still require evidence links and review.</figcaption></figure>

What is a building entity?

A useful building entity separates persistent identity from changing observations. It may include:

  • a stable identifier that survives new surveys and model versions;
  • one or more geometries with their coordinate systems and accuracy;
  • parcel, address, project and site relationships;
  • links to roofs, facades, floors, spaces and equipment;
  • dated captures and AI observations;
  • status, provenance and human-review state;
  • events such as construction, extension, demolition or subdivision.

The entity is therefore not a single mesh, polygon or BIM object. It is the record that connects those representations to the same real-world subject.

<figure><img loading="lazy" decoding="async" width="800" height="972" src="/media/insights/roof-surface-entities.webp" alt="Numbered partition of a complex roof into candidate surfaces and junction lines"><figcaption>Roof-surface partition from a Skylens experiment. The numbers identify review candidates; their junctions and ownership still require validation before they become structural entities.</figcaption></figure>

From geometry to identity

Entity matching asks whether a new observation belongs to an existing object. Location and overlap are important, but they are not enough. A robust process can also consider parcel membership, footprint similarity, roof structure, height, address, temporal continuity and neighboring objects.

The result should be a proposal with an explicit confidence level: match an existing entity, create a candidate entity, split one entity into several, merge candidates, or send the case for review. The system must preserve the observation even when the identity decision remains unresolved.

This distinction supports a simple rule: points and pixels are evidence; entities represent structure; relationships represent topology.

Time, versions and events

Reality-capture data is inherently temporal. Each flight, scan or panorama documents a particular state of the site. The entity should remain stable while its representations receive capture dates, valid-time ranges and processing versions.

Not every geometric difference is a physical change. A new outline may result from a different flight altitude, image quality, reconstruction setting or occlusion. A verified change event requires traceable evidence and, where the consequence matters, human approval.

This model makes it possible to ask: What did the system observe? What changed in the source data? What was accepted as a real-world change? Who approved it, and when?

Evidence before decisions

Every proposed attribute should point back to evidence: source image, point subset, model version, processing run and relevant quality metrics. The system should retain uncertainty instead of converting it silently into false precision.

For example, an AI mask can support a candidate roof surface. Multi-view imagery and registered point-cloud neighborhoods may strengthen it. A survey point may constrain position. A human reviewer can then approve, reject or refine the proposal. The approved entity remains connected to all of those inputs.

The working pipeline

  1. Preserve the original capture and its metadata.
  2. Register imagery, point clouds and models in a known spatial reference.
  3. Run detection or segmentation and retain the model version.
  4. Project observations into the shared spatial frame.
  5. Compare them with existing entities and relationships.
  6. Create a proposed match, new entity or review task.
  7. Validate geometry, confidence and provenance.
  8. Publish an approved version without deleting earlier evidence.

In Skylens Viewer, this approach supports navigation from a site or building to the captures, observations and tasks that explain its current state. It also creates a safer basis for search and AI assistance because the system can distinguish facts, proposals and unknowns.

Edge cases

Connected buildings may appear as one mask. One building may cross parcel boundaries. A temporary canopy can resemble an extension. A tree can obscure a roof edge. A construction site may contain an incomplete structure whose identity is known even though its geometry changes every month.

These are not reasons to discard automation. They are reasons to model ambiguity explicitly. Candidate states, review queues and provenance allow the workflow to move forward without pretending that every prediction is final.

Frequently asked questions

Can an AI mask be used directly for measurement?

Only when the source, scale, spatial reference, resolution and accuracy have been validated for that measurement. A visually plausible mask is not automatically survey-grade geometry.

What is the difference between a building entity and a 3D model?

A 3D model is one representation at a particular time and level of detail. The entity is the persistent record that can link several models, footprints, captures, documents and events to the same building.

Does every detection require human approval?

Not every low-risk observation needs the same review. Approval requirements should follow confidence, consequence and workflow rules. Identity changes and material geometry changes deserve stronger controls.

How should a team begin building persistent entities?

Start with stable identifiers, source provenance and a small set of clear relationships. Keep uncertain matches as candidates, test them on real projects and expand the ontology only when the workflow requires it.

What this means for Skylens projects

The goal is not to make AI produce a convincing colored layer. It is to build a reliable chain from observation to identity, version, evidence and action. That chain allows spatial data to remain understandable long after the original capture or model run.

Continue with 3D mapping and spatial models, or contact Skylens to discuss an evidence-based reality-capture workflow.