Digital Twin Ontologies: Brick Schema, RDF and SHACL for Semantic Building and Asset Models

Digital Twin Ontologies: Brick Schema, RDF and SHACL for Semantic Building and Asset Models

Digital Twin Ontology: Brick Schema, RDF and SHACL for Semantic Models

A building management system will happily tell you that point AI-7 on device 2012 reads 21.4. It will not tell you that this number is the air temperature of a particular room, that the room is served by a particular variable air volume box, that the box is fed by a particular air handler, or that the value is in degrees Celsius. Every analytics vendor then spends weeks reverse-engineering those facts from point names that were typed by hand in 2009. A digital twin ontology fixes this at the root: it gives every asset, space, point and relationship a machine-readable meaning, so software can ask questions instead of guessing.

This matters now because the work has moved from pilots to portfolios. A twin that is hand-wired to one site does not scale to fifty. The combination that has become the practical default in buildings is Brick Schema for the vocabulary, RDF for the data model, SPARQL for questions and SHACL for quality gates.

You will leave with a working mental model of each layer, runnable Python that builds a small building graph, queries it and validates it, a method for mapping BACnet points, and a governance checklist.

What this covers: the semantic stack and how Brick relates to RDF, OWL and SHACL; triples and relationships; SPARQL over a building model; SHACL shapes; BACnet and Modbus mapping; tooling; pitfalls; and a decision matrix against Haystack, ASHRAE 223P, DTDL and AAS.

Context and Background

Buildings and plants are full of data and almost empty of meaning. A typical campus has controllers from several vendors, each exposing points through BACnet or Modbus with naming conventions that reflect an installer’s habits rather than any shared standard. The names are strings. The relationships between the things the strings describe live in drawings, in a facilities engineer’s head, and sometimes in a spreadsheet nobody has opened since commissioning.

The research community framed this as a metadata problem. The early Brick work, presented at the BuildSys conference in 2016 under the title “Brick: Towards a Unified Metadata Schema for Buildings”, argued that building applications were not portable because each deployment needed bespoke metadata mapping. Several efforts have attacked the same problem from different angles. Project Haystack uses a tag-based model: an entity is described by marker tags such as site or ahu and value tags such as dis or area, and a separate set of definitions says what combinations of tags mean. RealEstateCore, developed in the Nordic real estate sector, offers an OWL ontology for property, spaces and building components. The W3C SOSA and SSN vocabulary, a Recommendation dated 19 October 2017, models sensors, observations, observable properties and features of interest in a domain-neutral way.

Brick sits in the middle of these. The project describes itself as an open-source effort to standardise semantic descriptions of the physical, logical and virtual assets in buildings and the relationships between them. It has three parts: an extensible dictionary of terms, a set of relationships, and a flexible data model. The repository is BSD-3-Clause licensed and follows semantic versioning, with minor releases targeted roughly every six months. When I checked the project releases page for this article, the latest stable release was 1.4.4, with a 1.5.0 release candidate and nightly builds ahead of it. Treat those version numbers as a snapshot and re-check before you pin anything.

Standards bodies are converging on the same technical substrate. ASHRAE’s proposed Standard 223P, titled “Semantic Data Model for Analytics and Automation Applications in Buildings” in its public review drafts, defines its model in RDF with SHACL shapes. NIST, a participant in the effort, stated that it would be published in fiscal year 2026; I could not confirm final publication from a primary source, so treat its status as something to verify. The 2018 announcement of collaboration between the ASHRAE BACnet committee, Project Haystack and Brick shows how long this convergence has taken.

The same questions arise outside buildings. Industrial teams building asset administration shells or Azure digital twin models face identical issues of meaning, units and relationships. Our comparison of AAS, DTDL and OPC UA information models covers that landscape, and the Brick approach here is a useful contrast: a graph-first model with open-world semantics rather than a typed-object model. Hold that distinction, because it explains most of what follows. For the underlying data model, the W3C RDF 1.1 specifications are the primary reference, and the Brick documentation at docs.brickschema.org is the primary reference for the building vocabulary.

The Semantic Stack: How Brick, RDF, OWL, SHACL and SPARQL Fit Together

Direct answer: A digital twin ontology is a formal vocabulary of classes and relationships plus the rules for using it. In the Brick stack, RDF stores facts as subject-predicate-object triples, Brick supplies the class and relationship vocabulary as an OWL ontology, SHACL validates that a model is complete and well-formed, and SPARQL queries it. Each layer does one job.

Digital twin ontology pipeline from BACnet point lists through Brick RDF graph and SHACL validation to SPARQL applications

Figure 1: The pipeline from raw point lists to a validated Brick graph that applications query with SPARQL.

The figure shows the part of the pipeline that most projects skip. Raw point exports and design models flow through mapping rules into an RDF graph, but the graph is not published until a validation gate passes. Violations flow back to whoever produced the data, not forward to whoever consumes it. That single design decision, putting the gate before the store, is what separates a semantic model that stays trustworthy from one that rots within a year.

RDF is the data model, not the schema

RDF represents every fact as a triple: a subject, a predicate and an object. Subjects and predicates are identified by IRIs, which are globally unique names; objects are either IRIs or literals such as numbers and strings. A set of triples is a graph, because the object of one triple can be the subject of another. RDF 1.1 is the W3C Recommendation that defines the abstract model, and it can be written in several syntaxes, of which Turtle is the most readable.

Two properties of this model matter for digital twins. First, there is no fixed table schema. You add a new kind of fact by writing a new triple, which is why a graph copes with the long tail of equipment that every real building contains. Second, identifiers are global. When two graphs from two vendors mention the same IRI, they are talking about the same thing, and merging them is a set union. Tabular and JSON-document models can approximate this, but they make you build the merge yourself.

Brick is the vocabulary, defined in OWL

Brick supplies the terms. Equipment, locations, points and quantities are classes arranged in a hierarchy; feeds, hasPoint, isPartOf, hasLocation and similar are properties that relate instances. The bundled copy I loaded for this article contains roughly 62,000 triples and declares 1,472 OWL classes, which includes aligned vocabularies, so it is a large dictionary rather than a handful of terms.

Because Brick is expressed in OWL, its classes carry subclass relationships and its relationships carry inverses. brick:feeds has brick:isFedBy as its inverse and brick:hasPoint has brick:isPointOf; I confirmed both in the ontology file. A Zone_Air_Temperature_Sensor is a kind of Air_Temperature_Sensor, which is a kind of Temperature_Sensor, which is a kind of Sensor, which is a kind of Point. That hierarchy is what lets a query ask for “any temperature sensor” and find the specific subtype that a particular building used.

Reasoning is what makes the hierarchy useful. A reasoner (the Python examples below use the OWL-RL profile through the brickschema package) materialises the implied triples, so a sensor declared as the most specific class is also found under every ancestor class. Brick also leans on aligned vocabularies. The repository vendors a pinned copy of RealEstateCore, and the 1.4.4 release deprecated many Brick quantity definitions in favour of QUDT equivalents for pressure, temperature and energy. The practical lesson: your model will mix namespaces, so write queries with prefixes and do not assume everything lives in brick:.

SHACL turns “valid RDF” into “usable model”

Here is the gap. RDF and OWL are open-world: the absence of a fact means “not stated”, not “false”. That is the right default for knowledge on the web and the wrong default for a commissioning checklist. If a VAV box has no zone temperature sensor, OWL does not complain. It assumes the sensor exists somewhere unstated.

SHACL, the Shapes Constraint Language, is the closed-world counterpart. It is a W3C Recommendation dated 20 July 2017. You write shapes that target a set of nodes, usually by class with sh:targetClass, and constrain the values reachable from them using property shapes with constraints such as sh:minCount and sh:class. A validator checks the data and returns a report. It is the same idea as a database schema constraint or a JSON Schema, applied to a graph.

The split of responsibilities is clean. OWL and the Brick ontology say what things mean. SHACL says what must be true of a particular model before you are willing to build applications on it. ASHRAE’s 223P work uses the same division, with its model defined in RDF and its rules expressed as shapes.

Where SOSA/SSN and Haystack fit

SOSA and SSN describe observations: a sensor made an observation of an observable property of a feature of interest. That is useful when your problem is provenance and measurement semantics, for example in environmental sensing or research data. Brick, by contrast, answers the facility question: which sensor belongs to which equipment, and which equipment serves which space. The two compose; a Brick point can reference the observable property it measures, and some projects publish time-series observations as SOSA while keeping the asset structure in Brick.

Haystack takes the tag route. Its documentation defines a mapping from Haystack defs to RDF so that the ontology can be used with standard semantic tools: marker tags become OWL classes, value tags become properties, and instance data exports as nodes. If your estate is already Haystack-tagged, that mapping is the bridge into the graph world, and you do not need to retag from scratch.

Building a Brick Model: A Runnable Walk-through

Theory is cheap, so here is a complete small model. All code in this article was run with Python 3, rdflib 7.6.0, pyshacl 0.40.1 and brickschema 0.8.0, installed into a clean virtual environment. The ontology graph bundled with that brickschema release reports version 1.5.0, a pre-release at the time of writing, so pin your own versions and re-run the examples against the release you actually deploy. Install with pip install rdflib pyshacl brickschema.

The model is one tower, one floor, two rooms, one air handler feeding two VAV boxes, and five points. Notice what the code encodes: containment with isPartOf, airflow with feeds, and instrumentation with hasPoint.

# model.py
from rdflib import Graph, Namespace, RDF

BRICK = Namespace("https://brickschema.org/schema/Brick#")
UNIT = Namespace("http://qudt.org/vocab/unit/")
BLDG = Namespace("urn:demo-campus#")

def build_model() -> Graph:
    g = Graph()
    g.bind("brick", BRICK); g.bind("unit", UNIT); g.bind("bldg", BLDG)

    def add(s, p, o):
        g.add((s, p, o))

    # Spatial hierarchy
    add(BLDG.Tower_A, RDF.type, BRICK.Building)
    add(BLDG.Floor_2, RDF.type, BRICK.Floor)
    add(BLDG.Floor_2, BRICK.isPartOf, BLDG.Tower_A)
    for r in ("Room_201", "Room_202"):
        add(BLDG[r], RDF.type, BRICK.Room)
        add(BLDG[r], BRICK.isPartOf, BLDG.Floor_2)
    add(BLDG.HVAC_Zone_2N, RDF.type, BRICK.HVAC_Zone)
    add(BLDG.HVAC_Zone_2N, BRICK.hasPart, BLDG.Room_201)
    add(BLDG.HVAC_Zone_2N, BRICK.hasPart, BLDG.Room_202)

    # Equipment: AHU feeds two VAV boxes, each feeding a room
    add(BLDG.AHU_1, RDF.type, BRICK.Air_Handling_Unit)
    add(BLDG.AHU_1, BRICK.hasLocation, BLDG.Tower_A)
    for v, room in (("VAV_201", "Room_201"), ("VAV_202", "Room_202")):
        add(BLDG[v], RDF.type, BRICK.Variable_Air_Volume_Box)
        add(BLDG.AHU_1, BRICK.feeds, BLDG[v])
        add(BLDG[v], BRICK.feeds, BLDG[room])

    # Points
    def point(name, cls, owner, unit):
        add(BLDG[name], RDF.type, cls)
        add(owner, BRICK.hasPoint, BLDG[name])
        add(BLDG[name], BRICK.hasUnit, unit)

    point("AHU1_SAT", BRICK.Supply_Air_Temperature_Sensor, BLDG.AHU_1, UNIT.DEG_C)
    point("AHU1_SF_Cmd", BRICK.Fan_Speed_Command, BLDG.AHU_1, UNIT.PERCENT)
    point("VAV201_ZT", BRICK.Zone_Air_Temperature_Sensor, BLDG.VAV_201, UNIT.DEG_C)
    point("VAV201_Damper", BRICK.Damper_Position_Command, BLDG.VAV_201, UNIT.PERCENT)
    point("VAV202_ZT", BRICK.Zone_Air_Temperature_Sensor, BLDG.VAV_202, UNIT.DEG_C)
    # VAV_202 deliberately has no damper command so validation can catch it
    return g

Three modelling choices deserve a comment. The room-to-floor link uses isPartOf and the HVAC zone uses hasPart because a zone is a logical grouping that aggregates spaces, while a floor is a physical container. The units are QUDT IRIs, the same units vocabulary that Brick is migrating its quantity definitions towards. And the points hang off the equipment they instrument, not off the room they happen to measure; the room is reached through the feeds path. Which of those is “right” depends on the questions you want to ask, and that is the real skill in ontology design.

Brick graph of an air handler feeding two VAV boxes and rooms with isPartOf containment and hasPoint links to sensors

Figure 2: The example model as a graph. Airflow edges use feeds, containment uses isPartOf, instrumentation uses hasPoint, and units attach to points.

The graph in Figure 2 has three overlapping topologies in one structure: the air path, the spatial hierarchy and the instrumentation. A table can hold any one of them naturally. A graph holds all three and lets a query walk between them, which is where the value is.

Querying the Model with SPARQL

SPARQL is the query language for RDF, standardised by W3C as SPARQL 1.1. A query is a graph pattern with variables; the engine returns every binding that makes the pattern true. Because Brick queries are pattern matches over relationships, they replace the brittle string logic (name LIKE '%ZN-T%') that every building analytics team has written at least once.

Run the model through the reasoner first. Without inference, a sensor declared as Zone_Air_Temperature_Sensor is invisible to a query for Temperature_Sensor.

# query.py
import brickschema
from model import build_model

g = brickschema.Graph(load_brick=True)   # Brick ontology plus helpers
g += build_model()
g.expand(profile="owlrl")                # materialise subclass and inverse triples

Q_ZONE_SENSORS = """
PREFIX brick: <https://brickschema.org/schema/Brick#>
SELECT ?vav ?sensor WHERE {
  ?vav a brick:Variable_Air_Volume_Box ;
       brick:hasPoint ?sensor .
  ?sensor a brick:Temperature_Sensor .
}"""
for r in g.query(Q_ZONE_SENSORS):
    print(r.vav.split("#")[1], r.sensor.split("#")[1])

Q_UPSTREAM = """
PREFIX brick: <https://brickschema.org/schema/Brick#>
SELECT ?src ?room WHERE {
  ?src a brick:Air_Handling_Unit .
  ?src brick:feeds+ ?room .
  ?room a brick:Room .
}"""
for r in g.query(Q_UPSTREAM):
    print(r.src.split("#")[1], "->", r.room.split("#")[1])

On the example model this prints the two VAV-to-sensor pairs, then AHU_1 -> Room_201 and AHU_1 -> Room_202 (row order is not guaranteed). The first query shows subclass inference at work: the sensors were declared with a more specific class, but the query asked for the general one. The second uses the property path brick:feeds+, which means “one or more feeds hops”. That is the capability a relational schema makes painful: you do not need to know how many hops separate the air handler from the room, and a model with an intermediate fan-powered box or a mixing section still answers correctly.

Real queries grow from these shapes. “Find every room whose zone temperature sensor is served by an air handler with a supply-air temperature sensor” is a three-line extension. “Which points lack a unit?” is the same query with FILTER NOT EXISTS. The reason analytics vendors adopt Brick is that these question templates become portable: the same SPARQL runs against any building whose model uses the same vocabulary, which is the portability that the early Brick research promised.

Performance and scale, honestly

A few thousand points fit comfortably in an in-memory rdflib graph. Beyond that, move to a triple store. Inference is the cost to watch: materialising OWL-RL closure can multiply triple counts, and feeds+ over a deep plant network is more expensive than a fixed-length pattern. I have not benchmarked stores for this article and will not quote throughput numbers I did not measure. Measure on your own model, keep the inferred graph as a rebuildable artefact rather than the source of truth, and separate the asserted model (what humans and tools wrote) from the inferred one.

SHACL Validation in Practice

Now the quality gate. The shapes below say three things: every VAV box exposes a zone temperature sensor and a damper command and feeds at least one room; every point declares exactly one unit; and every room belongs to exactly one floor. The VAV shape uses sh:qualifiedValueShape because a VAV has several hasPoint values and the constraint is about at least one of a given class, not about all of them.

# shapes.ttl
@prefix sh:    <http://www.w3.org/ns/shacl#> .
@prefix brick: <https://brickschema.org/schema/Brick#> .
@prefix ex:    <urn:shapes#> .

ex:VAVShape a sh:NodeShape ;
    sh:targetClass brick:Variable_Air_Volume_Box ;
    sh:property [
        sh:path brick:hasPoint ;
        sh:qualifiedValueShape [ sh:class brick:Zone_Air_Temperature_Sensor ] ;
        sh:qualifiedMinCount 1 ;
        sh:message "VAV box has no Zone_Air_Temperature_Sensor" ;
    ] ;
    sh:property [
        sh:path brick:hasPoint ;
        sh:qualifiedValueShape [ sh:class brick:Damper_Position_Command ] ;
        sh:qualifiedMinCount 1 ;
        sh:message "VAV box has no Damper_Position_Command" ;
    ] ;
    sh:property [
        sh:path brick:feeds ;
        sh:minCount 1 ;
        sh:class brick:Room ;
        sh:message "VAV box must feed at least one Room" ;
    ] .

ex:PointUnitShape a sh:NodeShape ;
    sh:targetClass brick:Point ;
    sh:property [
        sh:path brick:hasUnit ;
        sh:minCount 1 ; sh:maxCount 1 ;
        sh:message "Point must declare exactly one brick:hasUnit" ;
    ] .

ex:RoomShape a sh:NodeShape ;
    sh:targetClass brick:Room ;
    sh:property [
        sh:path brick:isPartOf ;
        sh:class brick:Floor ;
        sh:minCount 1 ; sh:maxCount 1 ;
        sh:message "Room must be part of exactly one Floor" ;
    ] .

The Python driver passes the Brick ontology as ont_graph and turns on RDFS inference. That is essential for the PointUnitShape: brick:Point is the target class, but the data declares only specific subclasses. pySHACL’s documentation describes the ontology graph as the place to supply the class hierarchy it needs for such targets.

# validate.py
import brickschema
from pyshacl import validate
from rdflib import Graph
from model import build_model

brick = brickschema.Graph(load_brick=True)       # ontology graph for subclass reasoning
data = build_model()
shapes = Graph().parse("shapes.ttl", format="turtle")

conforms, report_graph, report_text = validate(
    data_graph=data,
    shacl_graph=shapes,
    ont_graph=brick,
    inference="rdfs",
    abort_on_first=False,
)
print("conforms:", conforms)
print(report_text)

When run, the report contains exactly one violation: focus node VAV_202, message “VAV box has no Damper_Position_Command”. That is the planted defect, and it is the kind a human reviewer misses in a 400-box building. The sh:message strings matter more than they look. A report that says QualifiedMinCountConstraintComponent is for the validator author; a report that says “VAV box has no Damper_Position_Command” can be forwarded to an installer.

Sequence diagram of SHACL validation of a Brick model with ontology subclass closure and a CI pass or fail report

Figure 3: How a validation run proceeds. The engine loads the ontology for subclass closure, selects focus nodes by target class and evaluates property shapes before returning a report to the CI gate.

As Figure 3 shows, the engine does not search the whole graph for problems. It selects focus nodes from each shape’s targets, then evaluates constraints against those nodes. This has a consequence that bites later: a node that is not targeted by any shape is never checked, so coverage of your shapes defines coverage of your quality gate.

The silent typo problem

A SHACL gate that checks structure will not catch a misspelled class name. While writing this article I first used brick:Supply_Fan_Speed_Command for the air handler fan command, and the build and every query ran without complaint. The class is not defined in the ontology I loaded; the defined class is Fan_Speed_Command. Brick’s own embedded shapes did flag the point as “does not have class brick:Point”, but only because I had asked for that check explicitly. Plain RDF is happy to mint a new class from any IRI.

The remedy is a vocabulary lint run alongside shape validation: collect every term in the Brick namespace used as an instance type, and subtract the set of classes the ontology defines.

# lint.py
import brickschema
from rdflib import RDF, OWL
from model import build_model, BRICK

def unknown_brick_terms(data):
    onto = brickschema.Graph(load_brick=True)
    defined = set(onto.subjects(RDF.type, OWL.Class))
    used = {o for o in data.objects(None, RDF.type) if str(o).startswith(str(BRICK))}
    return sorted(str(u).split("#")[1] for u in used - defined)

print(unknown_brick_terms(build_model()))   # [] when every class is defined

If you deliberately use a typo’d class on a node, this prints it. It takes minutes to write and prevents the most common class of semantic defect I know of in hand-built models. The brickschema package also exposes a validate method that checks a graph against the shapes embedded with Brick; in my run, without the QUDT vocabulary loaded, it also reported items inside the ontology itself, so load QUDT or filter the report to your own namespace before wiring it into CI.

Mapping BACnet and Modbus Points into the Model

The graph is only as good as the link back to the real controller. The data you want to query lives in a time-series store or directly on a field bus, so each Brick point needs an external reference that tells a driver how to read it. Brick defines a reference vocabulary for this under https://brickschema.org/schema/Brick/ref#, and the brickschema Python package exposes the BACnet namespace http://data.ashrae.org/bacnet/2020# with properties such as object-identifier, object-type and device-instance. I checked those names in the installed package source rather than assuming them.

Mapping has two separate problems that teams constantly conflate. Classification answers “what kind of thing is this point?”, for example that VAV201.ZN-T is a zone air temperature sensor. Association answers “which equipment and space does it belong to?”. Classification can be partly automated from names and BACnet object types, because an analog input in degrees Celsius is probably a sensor. Association needs a source of truth about topology, which is usually a drawing, a BIM export or a person.

Here is a deliberately small rule-based classifier. It uses suffix rules, a unit table, and an explicit list of unmapped rows, because the rows a classifier cannot handle are the most valuable output it produces.

# mapping.py
import csv, io
from rdflib import Graph, Namespace, Literal, RDF
from rdflib.namespace import XSD

BRICK = Namespace("https://brickschema.org/schema/Brick#")
UNIT = Namespace("http://qudt.org/vocab/unit/")
BLDG = Namespace("urn:demo-campus#")
REF = Namespace("https://brickschema.org/schema/Brick/ref#")
BACNET = Namespace("http://data.ashrae.org/bacnet/2020#")

RAW = """object_name,object_id,device,units
AHU1.SAT,analog-input:1,1001,degC
AHU1.SF.CMD,analog-output:3,1001,percent
VAV201.ZN-T,analog-input:7,2012,degC
VAV201.DMPR,analog-output:2,2012,percent
VAV999.ZN-T,analog-input:7,2999,degF
AHU1.XYZ,analog-value:9,1001,none
"""

RULES = [
    ("SAT",    BRICK.Supply_Air_Temperature_Sensor),
    ("SF.CMD", BRICK.Fan_Speed_Command),
    ("ZN-T",   BRICK.Zone_Air_Temperature_Sensor),
    ("DMPR",   BRICK.Damper_Position_Command),
]
UNITS = {"degC": UNIT.DEG_C, "degF": UNIT.DEG_F, "percent": UNIT.PERCENT}

def equipment_id(prefix: str) -> str:
    return prefix.replace("AHU1", "AHU_1").replace("VAV", "VAV_")

def map_points(raw: str):
    g, unmapped = Graph(), []
    for row in csv.DictReader(io.StringIO(raw)):
        prefix, _, suffix = row["object_name"].partition(".")
        cls = next((c for s, c in RULES if s == suffix), None)
        if cls is None or row["units"] not in UNITS:
            unmapped.append(row["object_name"])
            continue
        pt = BLDG[row["object_name"].replace(".", "_").replace("-", "_")]
        eq = BLDG[equipment_id(prefix)]
        g.add((pt, RDF.type, cls))
        g.add((eq, BRICK.hasPoint, pt))
        g.add((pt, BRICK.hasUnit, UNITS[row["units"]]))
        obj_type, _, inst = row["object_id"].partition(":")
        ref = BLDG[f"ref_{pt.split('#')[1]}"]
        g.add((pt, REF.hasExternalReference, ref))
        g.add((ref, RDF.type, REF.BACnetReference))
        g.add((ref, BACNET["object-type"], Literal(obj_type)))
        g.add((ref, BACNET["object-identifier"], Literal(int(inst), datatype=XSD.integer)))
        g.add((ref, BACNET["device-instance"], Literal(int(row["device"]), datatype=XSD.integer)))
    return g, unmapped

if __name__ == "__main__":
    g, unmapped = map_points(RAW)
    print(len(g), "triples; unmapped:", unmapped)

The last row, AHU1.XYZ with no unit, lands in the unmapped list, which is the intended behaviour. A classifier that guessed would have produced a plausible-looking and wrong point. Note also VAV999.ZN-T in degrees Fahrenheit: the mapping happily records unit:DEG_F, and it is the equipment association step, not this script, that must notice that no VAV_999 exists in the topology. The orphan check belongs in SHACL, not in the mapper.

Modbus and other protocols

The same pattern generalises. BACnet is convenient because its object model (device, object type, instance, present value) is rich enough to carry reference data and the ASHRAE namespace is standardised. Modbus is poorer: a register address, a function code and a scaling factor carry no semantics, and the scaling convention is vendor-specific. For Modbus you typically keep a device profile (a table from register to quantity, unit and scale) under version control, then generate Brick points and an external reference carrying the address. The profile is the real asset; the RDF is derived from it.

The same applies to OPC UA and MQTT-fed plant data. In industrial settings the point reference is a node ID or a topic rather than a BACnet object, and the semantic layer sits above the protocol. If your estate also includes asset administration shells, the question of which unit vocabulary to anchor on arises immediately; our note on AAS units of measurement, IDTA 01003-B versus IEC 61360 shows the same unit problem from the industrial side. Brick’s move to QUDT and the IEC 61360 dictionary approach solve the same problem with different catalogues.

Tooling and Governance: Keeping the Model Alive

Building the first model is the easy part. The hard part is that buildings change: a tenant fit-out moves a wall, a controller is replaced, a point is renamed during a retrofit. A model that is correct at handover and unmaintained is a liability, because applications trust it.

Tooling

For modelling and validation, the practical open-source baseline is Python: rdflib for graphs, pySHACL for validation, and brickschema for the Brick ontology, reasoning and helpers. NREL’s BuildingMOTIF (Building Metadata OnTology Interoperability Framework) packages the workflow at a higher level. Its repository describes it as a toolset for creating, storing, visualising and validating building metadata, with an SDK that hides RDF graph handling, SHACL validation and schema translation, plus connectors for existing metadata sources. Its documentation is at buildingmotif.readthedocs.io. I am reporting its self-description; I did not exercise it for this article. Its README lists Brick, Haystack and ASHRAE 223P as supported or planned targets, so check which are current for your version. The Brick project also hosts a web explorer for the ontology at ontology.brickschema.org. I could not find a “Brick Studio” product I could verify, so I do not recommend or describe one.

For 223P, the Open223 site is a collection of permissively licensed tools that it says are not developed or endorsed by ASHRAE: ontology downloads in Turtle, documentation, a SPARQL query page backed by Oxigraph, and example models. It labels 223P a proposed standard and warns that its hosted services may not reflect the latest version, which is exactly the caution an engineer should apply.

The governance loop

Governance sounds bureaucratic, so make it mechanical. The model lives in Git as Turtle or as the templates that generate it. Every change is a pull request. CI runs the vocabulary lint and the SHACL shapes, and blocks the merge on failure. Releases get tags so applications can pin a model version. After publication, monitoring watches for drift: points that stopped reporting, equipment with no points, points with no equipment.

Governance loop for a digital twin ontology from point export through classification, review, SHACL validation, Git versioning and drift monitoring

Figure 4: A governance loop. Low-confidence classifications go to human review; validation failures return to review; only passing models are versioned and published, and monitoring feeds change requests back in.

The loop in Figure 4 has two feedback edges, and both matter. Validation failure returns work to a human, not to the classifier, because a failing shape usually means missing knowledge rather than a wrong rule. Monitoring returns change requests to the start, because the model is a living product with an owner, not a deliverable.

Name that owner explicitly. Assign someone the right to say yes or no to new terms, extension classes and shape changes. Without a named owner, shape libraries fork per project and the portability that justified the whole exercise disappears.

Naming and identifiers

Decide the IRI scheme on day one. A pattern such as https://example.com/site/<site-id>/<equipment-id> costs nothing now and avoids a migration later. IRIs should be stable even when the human-readable label changes; store the label as rdfs:label and never encode meaning that can change (a tenant name, a floor number after a renumbering) in the identifier. The example in this article uses urn:demo-campus# for brevity, which is acceptable for a demo and a mistake for production, since a urn: namespace cannot be dereferenced or governed.

Brick, Haystack, 223P, DTDL and AAS: Choosing a Model

No single model wins everywhere, because they were designed around different centres of gravity. Brick and 223P are graph-first and building-centric. Haystack is tag-first and pragmatic. DTDL and the asset administration shell come from the cloud-IoT and industrial-manufacturing worlds, where the unit of modelling is a typed device or a product instance rather than a facility graph.

Compare DTDL first. The Azure Digital Twins model language describes twin types with properties, telemetry, commands and relationships, and instances are checked against those types. It is closed and typed, which is excellent for generating code and APIs, and weaker for open-ended semantic reasoning over a vocabulary of thousands of equipment types. The AAS metamodel organises an asset into submodels with semantic identifiers that point to external dictionaries; our write-up of the AAS metamodel 3.2 and the IDTA 26-01 migration covers how that has evolved. AAS is strong when the question is “what does this manufactured product expose, in a form a supply chain can exchange”, and less natural when the question is “which rooms does this air handler serve”.

The honest answer for many sites is a combination: a Brick or 223P graph for the facility and system topology, with device-level models (AAS submodels or DTDL interfaces) for the equipment itself, linked by identifiers. The linking is where projects stall, because it needs a stable identifier shared by both worlds.

Dimension Brick Haystack ASHRAE 223P DTDL AAS
Core idea OWL ontology plus relations in RDF Marker and value tags plus defs RDF model with shapes, proposed standard Typed interfaces with relationships Submodels with semantic IDs
Best at Facility topology, points, portable analytics Fast tagging of existing point lists Rich connections and system modelling for controls Cloud twin APIs and codegen Product and asset data exchange
Validation SHACL, user-defined Tag rules and conventions SHACL shapes in the model Parser and type checks Metamodel and submodel templates
Open-world reasoning Yes, with OWL Limited Yes No Limited
Maturity signal Stable 1.4.x, 1.5 in progress Established, widely deployed Public review drafts, publication status to verify Vendor-driven specification IDTA-governed releases
Main risk Terminology sprawl, modelling skill Ambiguity across projects Complexity, tooling still maturing Platform coupling Building-topology gaps

Read the table as a map of risks, not a ranking. Several cells, especially maturity, will change within the life of this article, so confirm them against the current release notes.

How to choose

Pick Brick when the main value is portable analytics and fault detection across buildings, your team can run Python and SPARQL, and you want one open standard with a large tool ecosystem. Pick Haystack when you have a large existing tagged estate and need results quickly, and plan to bridge to RDF later using its mapping. Watch 223P closely when your use case is control-system configuration and you need detailed connection modelling; verify its publication status, and treat early adoption as an engineering investment. Pick DTDL when your platform is Azure Digital Twins and your applications are written against typed APIs. Pick AAS when the assets are manufactured products exchanged between organisations.

Whichever you pick, separate three concerns in the architecture: the vocabulary, the data and the validation rules. Teams that fuse them into a single proprietary schema lose the ability to migrate.

A note on 223P and Brick

The 2018 announcement described a collaboration to bring Haystack tagging and Brick data-modelling concepts into the proposed 223P. The practical effect since is that 223P provides finer-grained modelling of connections (what carries what medium between which components) than Brick’s feeds. Open223 and the Brick community both publish material on alignment. I did not verify the status of a formal mapping between the two, so if your programme depends on one, check the current documentation of both projects before committing. A conservative stance: model in Brick today, keep your mapping rules and templates in version control, and treat a later migration to or alongside 223P as a project you can plan rather than a surprise.

Trade-offs, Gotchas, and What Goes Wrong

Ontologies do not fix bad data. If the point names are wrong, the model will faithfully encode the wrong thing. The semantic layer exposes errors that were previously hidden by vagueness, which is useful and also politically uncomfortable.

Over-modelling. Brick offers over a thousand classes, and the temptation is to use the most specific one for every point. A sensor classified as the exact subtype carries more information, but a wrong exact subtype is worse than a correct general one. Classify at the level you can defend and refine later; subclass inference means queries for the general class still work.

Open-world surprises. Teams new to OWL write a query for “VAVs with no damper command” and expect a reasoner to answer it. The closed-world answer comes from SHACL or from a FILTER NOT EXISTS query, not from OWL entailment. Know which layer answers which kind of question.

Inference blow-up and stale closures. Materialised inferred triples go stale when the asserted model changes. Rebuild them rather than editing them, and never treat the inferred graph as an editing target.

Version drift in the ontology itself. Brick evolves. The 1.4.4 release deprecated many quantity definitions in favour of QUDT, and the 1.5.0 release candidate notes mention deprecating brick:Collection in favour of rec:Collection and removing brick:connectedTo. In my check, the ontology bundled with brickschema 0.8.0 had no brick:connectedTo. A model that relies on a deprecated term keeps working until the day you upgrade the ontology. Pin the ontology version in the model’s metadata and run an upgrade rehearsal in CI.

Unit confusion. Recording unit:DEG_F is correct, but an application that assumes Celsius will produce nonsense. SHACL can require a unit; it cannot require that downstream code converts. Add a conversion layer, ideally using the QUDT conversion data rather than hand-written factors.

Skills and ownership. The most common failure is organisational: the model is built by a consultant, handed over, and never updated. The governance loop is cheap, but only if someone owns it.

Security and privacy. A building graph describes physical layout, equipment and occupancy-related sensors. Treat the triple store as sensitive infrastructure data, apply access control at the graph or named-graph level, and avoid putting personal data (such as an occupant name against a room) in the same graph as equipment topology.

Practical Recommendations

Start small and prove value on one question. Choose an analytics use case, such as detecting VAV boxes whose zone sensor is missing or out of range, and model only what that use case needs. Expand by use case, not by completeness.

Treat the ontology version as a dependency. Record the Brick version in your model metadata, pin the Python package versions, and re-run the whole pipeline when you upgrade. Re-verify the version facts in this article before relying on them, since the stable release, the release candidate and the proposed standards have all been moving.

Put the validation gate in CI before you put any consumer on the model. Include three kinds of check: vocabulary lint for undefined terms, structural shapes such as the ones above, and topology sanity checks such as orphan equipment, unreachable rooms and units. Write human-readable sh:message text for each, aimed at the person who will fix it.

Keep raw sources, mapping rules and generated RDF separate. The mapping rules are the intellectual property; the RDF is a build output.

  • Choose an IRI scheme and an owner before the first import.
  • Pin Brick, rdflib, pySHACL and brickschema versions; record them in the model.
  • Run a vocabulary lint, SHACL shapes and orphan checks in CI.
  • Keep unmapped points in a visible backlog, never guess.
  • Separate asserted from inferred triples.
  • Link device-level models (AAS or DTDL) by shared identifiers, not by duplication.
  • Review shapes quarterly alongside the ontology release notes.

Frequently Asked Questions

What is a digital twin ontology?

A digital twin ontology is a formal, machine-readable vocabulary of the classes of things in a twin, such as air handlers, rooms and sensors, together with the relationships allowed between them and the constraints that make a model valid. It lets software discover what a data point means and how it connects to equipment and spaces, instead of relying on naming conventions. In buildings the common choice is Brick Schema expressed in RDF and OWL.

What is the difference between an ontology and a knowledge graph?

An ontology defines the vocabulary and rules: classes, relationships and constraints. A knowledge graph is the data expressed with that vocabulary: your specific building, its equipment and its points. In the Brick stack, the Brick ontology is the shared schema layer and your RDF model of a particular site is the knowledge graph. Both are stored as RDF triples, which is why they combine so easily.

Do I need OWL reasoning, or is SHACL enough?

You usually need both, for different jobs. Reasoning materialises implied facts, so a query for temperature sensors finds the specific subtypes. SHACL checks that a model is complete and well-formed, which reasoning cannot do because OWL assumes missing facts are merely unstated. For small models you can validate with the ontology supplied as extra context and skip full materialisation, but queries over class hierarchies need some form of inference.

How does Brick Schema compare with Project Haystack?

Brick is an OWL ontology with explicit relationships between entities, validated with SHACL. Haystack describes entities with marker and value tags and a set of definitions that give tag combinations meaning. Haystack is quick to apply to existing point lists, while Brick gives stronger formal semantics and reasoning. Haystack publishes an RDF mapping, so a tagged estate can be bridged into the graph world.

Is ASHRAE 223P a published standard?

I could not confirm publication from a primary source. The Open223 site labels it a proposed standard, and a NIST page said it expected publication in fiscal year 2026. Public review drafts exist and describe a semantic data model for analytics and automation applications in buildings. Check the ASHRAE standards page for the current status before citing it as final in a specification or contract.

Can I use Brick outside buildings?

Brick is designed for buildings and their systems, so for factory lines or process plants you will find gaps in equipment classes. Its RDF-based approach is reusable, and you can extend the ontology with your own classes in a separate namespace and validate them with SHACL. For manufactured-asset data exchange, the asset administration shell is usually a better fit, and the two can be linked by shared identifiers.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *