Design Model Transformation Testing Datasets for Architecture Early Design

Published: 7 January 2026| Version 1 | DOI: 10.17632/fst2mw9jbx.1
Contributors:
Jun Xiao,
,
,
,
,

Description

This repository provides benchmark datasets for validating the performance of the proposed model-transformation framework in early-stage architectural design. In early design practice, models typically include only primary structural elements and apertures, with implicit but well-defined spatial organization and minimal semantic annotation. These characteristics make early design models highly flexible, yet challenging for automated interpretation and downstream analysis. The transformation evaluated using these datasets consists of two sequential steps. First, unstructured boundary-representation (B-rep) design models are transformed into ontology-based Knowledge Graph (KG) or Labeled Property Graph (LPG) representations, where building elements, spaces, and their topological relationships are explicitly extracted and encoded. Second, the ontology models are further transformed into Building Energy Models (BEMs), enabling the derivation of performance-related metrics. Together, these transformations enable AI systems to explicitly interpret spatial topology in early design and to access performance indicators essential for performance-driven automated design, optimization, and decision-making. Two subsets are provided to support both theoretical and practical validation. Dataset 1 consists of procedurally generated, geometry-pure models with limited redundant elements, formatted in B-rep, LPG, and EnergyPlus IDF representations, and is intended for controlled validation of transformation correctness. Dataset 3 contains real-world building models characterized by complex spatial relationships and numerous redundant or non-spatial elements, designed to evaluate the robustness of the cleansing and topology-extraction modules under realistic conditions. The quantitative metrics defined in the manuscript should be applied to assess transformation accuracy, robustness, and efficiency across these datasets.

Files

Steps to reproduce

The provided Dataset 1 includes 100 generated model with 4 files for each case: in.geo as the B-rep geometries; edges.json and node.json as the NG of topology; in.idf as the energy model. Users can create more such models with the grasshopper script script/EvomassGenerate.gh. Besides, the *.geo file can be visualized by the grasshopper script located at script/VisualizeGeoInRhino.gh. A NG transformation align the json file created on grasshopper is provided via script/testTopology.py. In the open dataset, each case includes four files. The *.skp file is the original sketchUp model which can also be processed by SEFAIRA. The *.geo file is the exported B-rep geometries from the sketchUp model. The *\_out.geo file is the remodeled geometries after the transformation processed by CCR method. And the *.xml is the topology which record the building area of each spaces.

Institutions

Categories

Architectural Design, Building Energy Analysis, Knowledge Graph

Funders

Licence