IoT Device Monitoring: Observability Architecture, Metrics and Fleet Management (2026)
Last Updated: October 2026
A server that crashes stops answering requests, and a dozen systems notice within seconds. A device that fails in the field usually does something quieter. The radio drops, the firmware wedges in a loop that still pets the watchdog, a sensor drifts, or an over-the-air update half-applies, and the broker keeps reporting a healthy connection. Meanwhile the dashboard shows green because the last good sample is still the most recent sample. IoT device monitoring is the discipline of making those quiet failures loud without drowning the on-call engineer in noise or the finance team in storage bills.
Two things changed the economics recently. OpenTelemetry’s wire protocol and Prometheus’s native OTLP receiver mean one pipeline can now carry device metrics, logs and traces without custom glue. And fleets have grown past the point where per-device alerting is workable, so the interesting design problems moved to fleet-level signals, twin drift and cardinality.
What this covers: a reference architecture that separates presence, telemetry and desired-state planes; how MQTT keep-alive and Last Will actually behave; what OpenTelemetry can and cannot do on constrained hardware; how to alert on service level objectives rather than raw thresholds; and how to budget cardinality and retention before they budget you.
What Changed for October 2026
This is a ground-up rewrite of the April 2026 version. If you read the earlier post, treat its specifics as superseded.
- Observability, not just metrics. The April version framed monitoring as a metrics pipeline into a time-series database. This version treats presence, telemetry and desired state as three separate planes with different failure semantics.
- OpenTelemetry is now a first-class path. Prometheus accepts OTLP natively when started with the
--web.enable-otlp-receiverflag, and OTLP over HTTP is a practical transport for gateways. The old claim that Prometheus is unsuitable for push-based IoT is no longer accurate for gateway-level data. - Corrections. Sparkplug is an Eclipse Foundation specification (published as ISO/IEC 20237), not a Traction or Infiswift one. The ISO/IEC 27001 Annex A.14 reference was removed because it does not describe monitoring controls in the current edition. The statement that Prometheus defaults to 1,000 series per metric was removed because the Prometheus documentation does not specify any such default.
- Unsourced figures removed. The earlier post quoted per-device-month prices, monthly infrastructure budgets for a 100,000-device fleet, and broker capacity ceilings with no traceable source. They are gone. Where this post uses numbers, they are either quoted from a linked standard or labelled illustrative arithmetic.
- New sections. Presence handling with grace periods, twin drift as a first-class metric, multi-window burn-rate alerting adapted for fleets, and a cardinality budget worked example.
Context and Background
Server monitoring grew up around a pull model. A scraper reaches each target, asks for current numbers, and treats a failed scrape as evidence the target is down. That assumption fails for devices. Most sit behind NAT or cellular carrier gateways, wake on a duty cycle to save power, and cannot be reached unless they initiate contact. Monitoring therefore inverts: devices push, and “no news” has to be interpreted rather than assumed to be bad news.
The incumbent stacks reflect that history. Cloud platforms such as AWS IoT Core and Azure IoT Hub bundle a broker, a registry, a device twin and basic metrics, and they are convenient up to the point where you need fleet-wide queries or cross-vendor devices. Open source stacks typically combine an MQTT broker, a collector, a time-series store such as Prometheus-compatible storage or InfluxDB, and Grafana. Industrial sites add Sparkplug on top of MQTT, which defines birth and death certificates for edge nodes; our Sparkplug B reference architecture covers that layer in detail.
The vocabulary matters because teams conflate three things. Monitoring is watching known signals against known thresholds. Observability is the ability to ask new questions of the data you already emit, which means rich, high-dimensional events rather than a fixed dashboard. Fleet management is the control loop that acts on what monitoring finds: rolling out firmware, rotating credentials, quarantining a bad cohort. Device management protocols such as LwM2M define that control surface, and the LwM2M device management architecture post explains how it pairs with telemetry.
The authoritative references for the transport layer are the OASIS MQTT 5.0 specification and the OpenTelemetry specification. MQTT 5.0 defines the Keep Alive, Will and session semantics that presence detection depends on. OpenTelemetry defines the data model and OTLP protocol that increasingly carry the telemetry. Both are cited again below, and links are collected under Further Reading.
One framing point that shapes everything else: device telemetry is cheap to produce and expensive to keep. A sensor that reports once a minute produces 1,440 samples a day. Multiply by ten metrics and a hundred thousand devices and you have 1.44 billion samples a day before a single label has been attached. The architecture below is mostly about deciding which of those samples deserve to exist, where they are summarised, and which questions they must answer.
A Three-Plane Reference Architecture for IoT Device Monitoring
Production fleets are easiest to reason about when monitoring is split into three planes: presence (is the device connected and alive), telemetry (what is it measuring about itself and the world), and desired state (what should it be doing and does it agree). Each plane has different delivery semantics, retention needs and alert logic, and mixing them is the root cause of most noisy or blind setups.

Figure 1: Reference architecture for IoT device monitoring. Devices publish through an edge agent to a broker; presence events, telemetry and twin state travel on separate paths and meet only in the alert rules.
The diagram shows a device with an edge agent or gateway, a clustered MQTT broker, and three consumers hanging off the broker. A presence service reads connect and disconnect events and Last Will messages. A telemetry collector accepts metrics, logs and events, applies filtering and aggregation, and writes to a metrics store and an event store. A twin store holds reported and desired state. All three feed an alert rule layer that is based on service level objectives, and only then reach on-call engineers with a runbook link.
Why presence is its own plane
Presence answers a binary question, and it is answered by the connection, not by the data. A device that is connected and publishing nothing is a different failure from a device that is disconnected. Treating “no samples in five minutes” as the offline signal conflates a slow publisher, a broker backlog and a dead radio. The broker knows the truth about the connection: it holds the TCP session, tracks the Keep Alive timer, and can publish a Will message when the session ends abnormally.
Presence should therefore be derived from broker lifecycle events and Will messages, and telemetry should be used only as a secondary witness. The practical benefit is latency: presence changes are visible within roughly one and a half Keep Alive intervals, regardless of how often the device publishes metrics.
Why telemetry needs a collector in the middle
Writing device samples straight from the broker into a database couples the database schema to firmware. A collector layer, whether an OpenTelemetry Collector, Telegraf, Vector or a stream processor, gives you one place to rename metrics, drop labels, aggregate per cohort, and enforce a cardinality budget. Our comparison of OpenTelemetry Collector, Vector and Fluent Bit goes through the trade-offs between those options. For an edge-focused look at where Prometheus and Loki fit versus a full OpenTelemetry pipeline, see the OpenTelemetry vs Prometheus and Loki ADR.
The collector is also where you absorb the burstiness of fleets. After a regional outage, tens of thousands of devices reconnect and flush buffered samples within a few minutes. A collector with bounded queues and backpressure protects the stores; a direct write path turns the recovery into a second outage.
Why desired state is a monitoring concern
Most monitoring discussions stop at “what did the device report.” Fleets also need “what did we ask it to do, and did it comply.” A device twin or shadow stores both a reported and a desired document, and the difference between them is a signal in its own right. A device can be perfectly healthy and still be running last month’s firmware because an update request never reached it. That is invisible in CPU and memory graphs and obvious in twin drift, which is why the architecture routes twin state into the alert layer rather than leaving it as a lookup table.
The trade-off is write amplification. Every reported-state change is a write to a document store, and chatty firmware can turn a twin into a hot spot. The answer is to keep twins for slow-changing state such as firmware version, configuration hash and mode, and leave fast-changing numbers in the telemetry plane.
The three planes in one table
| Plane | Source of truth | Delivery need | Typical retention | Primary alert |
|---|---|---|---|---|
| Presence | Broker session and Will | Prompt, deduplicated | Short, events only | Fleet fraction offline by cohort |
| Telemetry | Device counters and sensors | At least once, bounded loss | Tiered, days to years | SLO burn rate on freshness and quality |
| Desired state | Twin or shadow documents | Eventual, versioned | Current plus audit trail | Drift older than rollout window |
The sections that follow take each plane in turn, then show how to combine them into alerts that are worth waking someone for.
Presence: Keep Alive, Last Will and the Grace Period
The presence plane is simple in principle and full of edge cases in practice. This section walks through what MQTT guarantees, what it does not, and how to turn raw lifecycle events into an offline signal that does not flap.
What MQTT 5.0 actually guarantees
In the OASIS MQTT 5.0 specification (section 3.1.2.10), the client declares a Keep Alive interval, and if the server receives no control packet from the client within one and a half times that period, it must close the network connection. The client keeps the connection warm by sending PINGREQ packets when it has nothing else to send. This is the mechanism that detects half-open TCP sessions, where a device vanished without sending a FIN.
Two consequences follow. First, detection latency is bounded by Keep Alive, not by your publish interval. With a Keep Alive of 60 seconds, the worst case for the broker to notice a silent death is about 90 seconds, which is plain arithmetic from the 1.5 multiplier. Second, Keep Alive is a trade against battery. A cellular device that wakes every 60 seconds only to ping consumes more energy than one that pings every 15 minutes, and carriers often expire NAT mappings in a few minutes anyway, so very long intervals can cause the broker to believe a device is alive while the return path is already broken. Choosing Keep Alive is a negotiation between detection speed, radio energy and NAT behaviour, and the value should be set per device class rather than fleet-wide.
MQTT 5.0 also adds a Server Keep Alive property. If the broker returns it in CONNACK, the client must use that value instead of the one it sent. That lets operators cap the worst-case detection time centrally, which is useful when firmware in the field was built with an overly generous setting.
Last Will and Testament, and the Will Delay Interval
A client may register a Will message in its CONNECT packet. The broker publishes it if the connection ends without a clean DISCONNECT, for example on a Keep Alive timeout or a network error. The usual pattern is for the device to publish a retained “online” status on connect and register a retained “offline” Will on the same topic, so a late-subscribing consumer always sees the current state.
MQTT 5.0 added the Will Delay Interval (section 3.1.3.2.2): the server delays publishing the Will until that interval has passed or the session ends. This is exactly the right tool for flaky links. If a device drops and reconnects within the delay, the session resumes and the Will is never published, which removes a whole class of offline flaps caused by momentary signal loss. The related Session Expiry Interval controls how long the broker keeps session state after disconnect, with the maximum value 0xFFFFFFFF meaning the session never expires. Will Delay should be shorter than or equal to session expiry for the behaviour to be predictable.
Plan for what Will does not give you. It is not published on a clean disconnect, so a device that shuts down deliberately must announce that separately. It is best-effort, because if the broker itself fails the Will for in-flight sessions may be lost depending on persistence and clustering. And a Will is a single message with a single topic, so it cannot carry a rich explanation of why the device died.
Managed broker lifecycle events
Cloud brokers expose their own presence stream. AWS IoT Core, for instance, publishes events on topics such as $aws/events/presence/connected/<clientId> and $aws/events/presence/disconnected/<clientId>, and the disconnect payload includes a disconnectReason such as CLIENT_INITIATED_DISCONNECT, CONNECTION_LOST or MQTT_KEEP_ALIVE_TIMEOUT. That reason code is valuable: it separates a deliberate reboot from a network failure and from a duplicate client ID, which usually means two firmware images fighting over one identity.
The AWS documentation is explicit about a caveat that applies to many platforms. Lifecycle messages might arrive out of order, and you might receive duplicates. AWS recommends a wait state before acting on a disconnect, for example a short SQS delay queue followed by a check that the device is still offline. Without that, a disconnect followed by a quick reconnect can arrive as “connected, then disconnected” and leave your registry permanently wrong. Use the event timestamps and the session identifier to order events rather than trusting arrival order.

Figure 2: Presence detection sequence. The broker closes the silent session after one and a half Keep Alive periods, emits the disconnect and Will, and the presence service waits and re-checks before raising an alert.
The sequence in Figure 2 starts with a device connecting with a Will and a Keep Alive value. The broker emits a connected event, and the device publishes heartbeats and metrics. When the network drops silently, nothing arrives until the broker’s timer expires at one and a half Keep Alive periods. The broker then emits a disconnect event and the Will. The presence service does not alert immediately: it waits a grace period, queries the broker or registry to confirm the device has not reconnected, and only then raises an offline alert.
Heartbeat design: connected is not the same as healthy
A live TCP session proves the network stack and the MQTT client task are running. It does not prove the application is working. A common field failure is a firmware main loop that has deadlocked while a separate network thread keeps answering pings. The broker sees a perfect session; the sensor data stopped hours ago.
The remedy is an application-level heartbeat that exercises the real work path. Good heartbeats carry a monotonic counter incremented by the main loop, uptime, a reset-reason code and the firmware version, and they are published on a schedule driven by the application task rather than by the network library. Alert on counter stagnation, not on message arrival, because a replayed or buffered message can arrive without the loop being healthy. Including the reset reason lets you detect boot loops: a device that reports “brownout reset” five times in an hour is a power problem, not a software problem.
Heartbeat frequency is a cost decision. Illustratively, a 60-second heartbeat of about 100 bytes of payload amounts to roughly 144 KB a day per device before protocol overhead, which is negligible on Ethernet and meaningful on a metered cellular plan with thousands of devices. Many fleets choose a slow heartbeat, with an immediate event on state change, and let the presence plane cover fast failure detection.
Turning events into an offline signal
A robust presence service keeps a small state machine per device: connected, suspect, offline. A disconnect moves the device to suspect and starts a timer equal to your grace period. A reconnect inside the grace period returns it to connected with no alert. Expiry moves it to offline and emits the event that alert rules consume. Record the disconnect reason with each transition, because cohorts with the same reason are the first clue in diagnosing outages.
Choose the grace period from observed behaviour, not intuition. Measure the distribution of disconnect-to-reconnect gaps over a few weeks. If most reconnects occur within two minutes and a long tail follows, a five-minute grace removes most transients while still giving timely signal. Devices on duty-cycled radios need a different treatment altogether: for them, “offline” should mean “missed N expected check-ins,” where N is derived from their declared schedule, which the twin can store.
Telemetry: What to Collect, and How OpenTelemetry Fits Constrained Devices
The telemetry plane carries numbers, logs and traces about the device and its work. The main design question is not which database to use but how much structure to put on the device versus at the edge.
A metrics hierarchy that survives scale
Organise metrics by who consumes them. Device-level signals describe the device itself: uptime, reset reason, free heap, radio signal strength, retry counts, queue depth of unsent samples, battery voltage and the age of the last successful cloud acknowledgement. Process-level signals describe the work: sensor read failures, calibration status, actuator error counts. Gateway-level signals describe aggregation points: connected children, ingest rate, buffer occupancy and upstream latency. Fleet-level signals are derived downstream: fraction reporting on time per cohort, firmware version distribution, reboot rate per thousand device-days.
The most useful single device metric is freshness, the age of the newest accepted sample. It collapses connectivity, firmware health and pipeline delay into one number you can write an SLO against. The second is the unsent-queue depth on the device, because a growing queue means the uplink is failing even when some samples still trickle through. The third is reset count with reset reason, which turns reliability from an anecdote into a rate.
Prefer counters and gauges that the device accumulates, and let the backend compute rates. A device that reports “reboots since birth” lets you recover from lost messages, because the next successful report restores the correct total. A device that reports “reboots in the last minute” loses information whenever a message drops. This is the same reason Prometheus prefers cumulative counters.
OpenTelemetry on constrained devices: where it fits and where it does not
OpenTelemetry defines a vendor-neutral data model for metrics, logs and traces, plus the OpenTelemetry Protocol (OTLP), which carries them over gRPC or HTTP using protobuf. The current specification status page lists OTLP 1.11.0 alongside specification 1.61.0 and semantic conventions 1.44.0 at the time of writing. The reference SDKs target servers and mobile apps. They assume a heap, threads and a TLS stack that a 256 KB microcontroller does not have.
The practical split is therefore to speak OpenTelemetry at the gateway and use something thinner on the device. A Cortex-M class sensor publishes a compact payload, such as CBOR or a small protobuf, over MQTT or CoAP. A gateway or edge collector decodes it, attaches resource attributes, and emits OTLP upstream. Devices that run Linux, such as industrial gateways and single-board computers, can run the OpenTelemetry Collector or a lighter agent directly and receive OTLP from local processes. Our Grafana Alloy and OpenTelemetry Collector tutorial shows a working gateway configuration.
This split has a real benefit: the device payload format can stay stable for the life of the hardware while the gateway absorbs changes in naming conventions and backends. It also has a cost. The gateway becomes a translation layer that someone must own, version and test, and a bug there corrupts data for every device behind it.
Semantic conventions need care. OpenTelemetry defines device resource attributes such as device.manufacturer, device.model.identifier and device.model.name, all currently at Development stability, and device.id, which is opt-in and carries an explicit privacy warning in the specification because it may be personally identifiable. For industrial fleets, a stable asset identifier is usually justified, but you should decide deliberately which identifier goes into telemetry labels and which stays in the registry. Because attribute names at Development stability may change, isolate them behind your collector’s transform stage rather than hard-coding them in firmware.
Prometheus, OTLP and the push question
The old advice was that Prometheus is pull-based and therefore awkward for push-heavy IoT. That is still true for scraping individual devices, which you should not do. But Prometheus 3.0 can accept OTLP directly when started with the --web.enable-otlp-receiver flag, which is disabled by default because Prometheus can run without authentication. Resource attributes can be promoted to labels through an otlp section in the configuration, and UTF-8 metric and label names are supported so that dotted OpenTelemetry names no longer need mangling.
Be aware of the limits. Delta temporality, which low-power devices often prefer because they can reset after each report, was described as experimental in the Prometheus and Grafana guidance, with the recommended workaround being the delta-to-cumulative processor in the OpenTelemetry Collector. Practically, convert at the gateway, keep cumulative values in the store, and do not let device-side resets leak into queries as negative rates. If you use other stores such as InfluxDB or a managed vendor backend, check their temporality handling before assuming parity.
Cardinality is the real cost driver
A time series is one unique combination of metric name and label values. The Prometheus naming guidance warns against storing high-cardinality dimensions such as user IDs or other unbounded sets of values in labels, because each distinct combination creates a new series. A device identifier is bounded, but with a large fleet it behaves like an unbounded one.
Worked example, illustrative only. Suppose 100,000 devices each export 12 metrics, and every series carries a device label. That is 1.2 million series before any other dimension. Add firmware version with 8 live values, and in the worst case where every combination occurs, you approach 9.6 million. Add hardware revision with 5 values and you reach 48 million. Real fleets are sparse, because most devices run one or two firmware versions, but the multiplicative risk is why dimension additions need review.
Three controls keep this under control. First, do not put device identity on every metric. Keep per-device series for a short list of diagnostic metrics, and publish cohort aggregates, such as counts and histograms by firmware and region, for everything else. Second, put high-cardinality attributes on logs, traces or the registry, where they cost storage per event rather than per series. Third, enforce a series budget in the collector with drop rules and alert on growth of active series, so a firmware release that adds a label fails visibly in staging.
Retention follows the same logic. Keep raw per-device data for the window in which you will actually debug, often days to a few weeks, then downsample to cohort-level rollups for trend analysis. Illustrative arithmetic: one sample per minute for 12 metrics across 100,000 devices is 1.44 million samples per minute, or about 2.07 billion per day. Whatever your store’s per-sample cost, that number shows why retention tiers are not optional.
Desired State: Twin Drift as a Monitoring Signal
Reported telemetry says what the device claims. A twin adds what you asked it to do. Together they let you measure convergence, which is the property fleet operations actually care about during rollouts, configuration changes and quarantines.
The AWS IoT Device Shadow document is a representative model. It has a state section with desired and reported subsections, a computed delta, per-attribute metadata timestamps, and a version number used for optimistic concurrency. The delta contains fields present in desired that differ from reported. Fields only in reported do not appear in the delta, and arrays are treated as atomic values and replaced entirely. Azure IoT Hub device twins and Eclipse Ditto follow a similar desired/reported pattern, and the Eclipse Ditto tutorial shows an open implementation.

Figure 3: Twin drift flow. A rollout system writes desired state, the device writes reported state, and the delta becomes both a command to the device on reconnect and a fleet-level drift metric.
Figure 3 shows the loop. A rollout system or operator writes desired state. The device reports its actual state. The twin computes the delta, pushes it to the device when it connects, and exposes it to monitoring as a drift metric: how many devices currently have a non-empty delta, and how old is the oldest. An alert fires when drift outlives the expected rollout window.
Why drift beats polling for rollouts
Consider a firmware rollout to 20,000 devices over 48 hours. Polling each device for its version is expensive and misses offline devices. Reading the twin gives an instant distribution: reported version counts, devices with an open delta, and the age of each delta. The decision to pause or continue a rollout can then be automated. Progressive delivery systems such as those described in the Argo Rollouts post for edge fleets apply the same principle: promote only if the health signals of the early cohort stay within bounds.
Twin failure modes
The versioning field exists for a reason. Two writers updating desired state without version checks can overwrite each other, and a device that applies a stale delta after a long offline period may roll back a newer configuration. Always include the expected version in writes that matter, and design devices to ignore a delta whose version is older than the one they have applied.
Atomic arrays surprise teams. If desired state contains a list of allowed endpoints and an operator changes one entry, the whole array is replaced; a device that merges instead of replaces will diverge. Document merge semantics per field and test them.
Finally, do not turn the twin into a telemetry store. Document sizes are limited, writes are rate-limited on managed services, and fast-changing values create churn in the metadata timestamps. Keep the twin for configuration, mode, versions and last-known state, and keep samples in the metrics store.
Alerting on Fleets: SLOs, Burn Rates and Cohorts
Alerting is where monitoring becomes operations. The failure patterns are familiar: thresholds copied from a datasheet, thousands of identical pages after a regional outage, and alerts that nobody trusts because half are false. The fix is to alert on symptoms that matter to the service and to treat the fleet statistically.
Define SLIs that describe the service, not the box
A service level indicator is a ratio of good events to valid events. For devices, useful SLIs include the fraction of expected reports that arrived within the allowed freshness window, the fraction of readings passing quality checks, and the fraction of commands acknowledged within a deadline. Pick a target per device class. A safety-relevant controller might target 99.9% on-time reporting, while a battery-powered environmental sensor might target 98%, accepting that a few missed check-ins are normal.
Denominators matter. A device that is intentionally powered off, in maintenance or on a declared sleep schedule should be excluded from the valid set, which is why registry and twin state belong in the SLI computation and why the twin should record expected schedules.
Burn-rate alerts, adapted from the SRE Workbook
The Google SRE Workbook describes multi-window, multi-burn-rate alerting. For a 99.9% SLO it suggests paging when budget burns at 14.4 times the sustainable rate over a one-hour window, confirmed by a five-minute short window, which consumes about 2% of the budget; paging at a burn rate of 6 over six hours with a 30-minute short window, consuming 5%; and opening a ticket at a burn rate of 1 over three days with a six-hour short window, consuming 10%. Both windows must exceed the threshold, so alerts fire only while budget is actively being consumed and clear soon after recovery.

Figure 4: Alert design flow. Burn rate decides between page and ticket, while correlated fleet-fraction signals decide between an incident and an individual device work queue.
Those parameters were written for request-serving services and the numbers should be treated as a starting point. Fleet SLIs usually move more slowly than request error rates, so longer windows are often right. The principle transfers cleanly: pair a long window that proves significance with a short window that proves it is still happening.
Cohort correlation: one incident, not ten thousand pages
As Figure 4 shows, fleet-fraction signals follow a second path. Group devices by the dimensions that explain shared failures: firmware version, hardware revision, site, carrier, gateway and certificate batch. When a large share of one cohort goes offline together, that is an incident with a hypothesis. When devices go offline independently, it is a maintenance queue handled in batches during business hours.
Implement this by computing the offline fraction per cohort and alerting on that ratio with a minimum cohort size so tiny groups do not page. Route one alert per cohort, with the common attribute in the title and a link to the list of affected devices. Individual device failures create tickets in a work queue sorted by customer or asset impact.
Alertmanager-style grouping, inhibition and routing implement this directly: group by cohort labels, inhibit per-device alerts when a cohort alert is firing, and route by severity. The mechanism is configuration, but the discipline is deciding which labels define a cohort.
Trade-offs, Gotchas, and What Goes Wrong
Fast detection versus false alarms. Short Keep Alive and short grace periods find dead devices quickly and also flap on every transient. Long ones are quiet and slow. There is no universal setting; measure your reconnect distribution and set per device class.
Edge aggregation versus forensic detail. Aggregating at the gateway cuts cost and cardinality, but an average can hide the one device that is failing. Keep a small always-on set of per-device diagnostics, and retain the ability to request detailed capture on demand through the twin.
Cloud twin convenience versus lock-in. Managed shadows and registries are easy to adopt, but their document formats, size limits and event semantics differ. If you may change vendors, keep your own canonical schema and treat the vendor twin as a transport.
Monitoring the monitor. A silent collector looks identical to a healthy fleet. Run a canary device that publishes a known heartbeat through the same path, and alert if the canary goes stale. Also alert on ingest rate falling sharply, since a pipeline stall otherwise appears as “no alerts.”
Security telemetry is not optional. Monitoring channels are an attack surface. Use per-device credentials, topic-level authorisation so a device can publish only its own topics, and certificate expiry as a monitored metric. A fleet-wide certificate expiry is a predictable outage that monitoring should have warned about weeks earlier. The device identity and attestation architecture post covers the identity side.
Anti-patterns worth naming: alerting on CPU percentage without a symptom, a per-device page for every offline event, using the twin as a time-series database, writing device ID into every metric label, and runbooks that nobody has tested.
Practical Recommendations
Start with presence and freshness, because they cover the most failures for the least complexity. Add the twin drift metric before your first large rollout, not after. Introduce OpenTelemetry at the gateway where it is cheap, and keep device payloads small and stable. Budget cardinality on paper before any dimension ships.
A short rollout order that works for most fleets:
- Set Keep Alive per device class, add a Will message, and use Will Delay Interval on unreliable links.
- Build the presence state machine with a measured grace period and disconnect-reason capture.
- Publish an application-level heartbeat with a main-loop counter, reset reason and firmware version.
- Define a freshness SLI and SLO per device class, excluding declared sleep and maintenance.
- Alert on cohort offline fraction and on burn rate, with page versus ticket separated.
- Track twin drift age and block rollout promotion when drift or reset rates exceed bounds.
- Add a canary device and pipeline-health alerts, then review series growth monthly.
For a first alert rule, a cohort-level freshness check in Prometheus-style syntax illustrates the shape. The metric names are examples, not a standard:
groups:
- name: fleet-presence
rules:
- alert: CohortStale
expr: |
(sum by (firmware, site) (device_report_age_seconds > 600))
/ sum by (firmware, site) (device_expected_reporting)
> 0.10
for: 10m
labels: {severity: page}
annotations:
summary: "More than 10 percent of {{ $labels.firmware }} at {{ $labels.site }} is stale"
If your devices are machines being watched for wear rather than connectivity, pair this with the approach in condition monitoring and machinery health architecture. And for the relationship between monitoring data and richer models, IoT versus digital twin explains where a twin goes beyond state mirroring.
Frequently Asked Questions
How do I detect an IoT device that is connected but not working?
Use an application-level heartbeat rather than relying on the MQTT session. Publish a counter incremented by the main work loop, plus reset reason and firmware version, and alert when the counter stops advancing even if messages still arrive. Pair it with a freshness SLI on real sensor data. A live connection only proves the network task is running, so you need evidence from the code path that does the actual job.
What is the difference between Keep Alive and Last Will in MQTT?
Keep Alive is the interval the client promises to communicate within. If the broker hears nothing for one and a half times that interval, it closes the connection. Last Will is a message registered at connect time that the broker publishes when the connection ends without a clean disconnect. Keep Alive detects the failure; Will announces it to subscribers. MQTT 5.0 adds Will Delay Interval so brief reconnects do not trigger the Will.
Can I use OpenTelemetry on a microcontroller?
Not the standard SDKs, which assume resources most microcontrollers lack. The common pattern is a compact payload from the device over MQTT or CoAP, with a gateway or collector converting it to OTLP. Devices running Linux can run the Collector or a lightweight agent. Keep naming and attribute mapping in the gateway so firmware does not need to change when conventions evolve.
How many metrics per device is reasonable?
There is no fixed number, because cost depends on series count and retention, not metric count alone. A practical approach is a small per-device diagnostic set, such as freshness, queue depth, reset count and signal strength, and cohort aggregates for everything else. Calculate series as devices times metrics times label combinations, and set a budget before shipping a new label.
Should alerts be per device or per fleet?
Both, with different routes. Fleet and cohort alerts, such as the fraction of a firmware version that is stale, represent incidents and can page. Per-device failures usually belong in a ticket queue worked in batches, ranked by asset or customer impact. Inhibit per-device alerts while a cohort alert is firing so a regional outage does not create thousands of notifications.
Is Prometheus suitable for IoT fleets?
Not for scraping individual devices, since most cannot be reached inbound. It is a reasonable store for gateway-level and cohort-level metrics, especially now that Prometheus 3.0 can receive OTLP when the OTLP receiver flag is enabled. Check temporality handling and your series budget. For very large per-device data, consider a store designed for high cardinality, or aggregate before storage.
Further Reading
Related posts on this site:
- MQTT Sparkplug B reference architecture for birth and death certificates in industrial deployments.
- Sparkplug B versus plain MQTT topics for deciding whether you need the extra structure.
- LwM2M IoT device management architecture for the control plane that acts on monitoring signals.
- OpenTelemetry Collector versus Vector versus Fluent Bit for the collector layer.
- Argo Rollouts for edge fleets for health-gated rollouts.
External references:
- OASIS MQTT Version 5.0, sections on Keep Alive, Will and session expiry.
- Google SRE Workbook, Alerting on SLOs for multi-window burn-rate alerting.
- AWS IoT lifecycle events and Device Shadow documents for managed presence and twin semantics.
- OpenTelemetry device semantic conventions and Prometheus naming practices.
By Riju — about
