SLO Error Budget for IoT Platforms: Burn-Rate Alerting with OpenSLO and Sloth
Most IoT platform teams alert on the wrong thing. They page when CPU on a broker node crosses 80 percent, or when a single device goes quiet, and they sleep through the slow leak where 3 percent of telemetry silently never reaches the digital twin. The platform is “up” by every infrastructure metric while the product, a trustworthy live model of physical assets, is degraded. A well-built SLO error budget fixes this by expressing reliability as a measurable promise to the consumer of the data and by paging only when that promise is being spent faster than the business can afford.
This matters in 2026 because IoT and digital-twin platforms now feed control loops, maintenance decisions and customer-facing dashboards, and the old “is the broker alive” health check says nothing about data quality. Stale or missing data is the failure mode, and it needs its own indicators.
You will leave with a concrete method: IoT-specific service level indicators (SLIs), the burn-rate arithmetic behind the 14.4x, 6x and 1x thresholds, an OpenSLO document, the Prometheus rules Sloth generates from it, and an error budget policy a team can actually sign.
What this covers: why IoT needs different SLIs, the reference architecture for measuring them, the burn-rate math, OpenSLO and Sloth in practice, low-traffic and fleet-skew gotchas, and a rollout checklist.
Context and Background
Service level objectives entered mainstream engineering practice through Google’s site reliability engineering (SRE) books. The vocabulary is small. A service level indicator (SLI) is a measured ratio of good events to valid events. A service level objective (SLO) is a target for that ratio over a window, such as 99.9 percent of requests over 30 days. The error budget is the complement: the 0.1 percent you are allowed to get wrong. An SLA, by contrast, is a contractual commitment with financial consequences, and it should always be looser than the internal SLO.
The reason the model works is that it converts an argument (“is the platform reliable enough?”) into arithmetic (“we have 31 minutes of budget left this month”). Product and engineering negotiate the target once, then let the numbers arbitrate release pace. The Google SRE Workbook chapter on alerting on SLOs is the canonical treatment, and the multiwindow, multi-burn-rate alert design it describes is what every modern tool implements.
Most published SLO guidance assumes a request-response service: an HTTP API where each request succeeds or fails within milliseconds. IoT platforms break that assumption in three ways. First, the interesting unit is not a request but a message or a reading that travels through several asynchronous hops: device, gateway, broker, stream processor, time-series database, twin model, dashboard. A failure at any hop is invisible to the device and often invisible to the user until they notice a flat line. Second, traffic is shaped by physics and schedules, not people, so volume can be extremely smooth (a fixed reporting interval) or extremely bursty (a plant restart reconnecting 40,000 devices). Third, devices go offline for legitimate reasons, so “no data” is ambiguous between an outage in your platform and a truck driving through a tunnel.
Tooling has also matured. OpenSLO gives a vendor-neutral YAML specification for SLOs, and Sloth turns SLO definitions into Prometheus recording and alerting rules that implement the Google multiwindow design. Sloth’s GitHub releases page lists v0.16.0 as the latest tag at the time of writing (dated 4 April 2026 in the release listing I fetched), and it supports its own default spec, Kubernetes custom resources and OpenSLO input. If you already run Prometheus, or a Prometheus-compatible store, you can have burn-rate alerting working in an afternoon.
If you are still deciding how telemetry about the platform itself should be collected, the trade-offs between agent-based and pipeline-based approaches are covered in OpenTelemetry versus Prometheus and Loki for edge observability, and the transport question is covered in OTLP versus Prometheus remote write. This post assumes the metrics already exist and focuses on what to measure and when to wake someone up.
The Reference Pattern: Four SLIs Along the Data Path
A good IoT platform SLO error budget is built from SLIs that measure the data product, not the servers: message loss, ingestion latency, data freshness, and device-to-dashboard latency. Each is a ratio of good events to valid events over a rolling 30-day window, measured at the consumer side of each hop, and each gets its own target and budget.

Figure 1: Four SLIs measured along the IoT data path. Each dotted line marks where the measurement is taken, not where a component lives.
The diagram shows one path and four measurement spans. The spans overlap deliberately. Message loss and ingestion latency localize problems to the front of the pipeline, freshness localizes them to the store and twin layer, and device-to-dashboard latency is the end-to-end number that the user actually experiences. When the end-to-end SLI burns, the inner SLIs tell you where.
SLI 1: Message loss (completeness)
Message loss is the ratio of messages accepted by the platform that reach durable storage, over messages that devices attempted to send. The numerator is easy: count rows or points written. The denominator is the hard part, because the platform cannot count what never arrived. There are three practical ways to build the denominator.
The first is device-side sequence numbers. If every device increments a counter per message, the ingest pipeline can compute gaps per device, and the sum of gaps plus received messages is the expected total. This is accurate but requires firmware cooperation, and gaps caused by a device that rebooted and reset its counter must be handled. The second is broker-side counters: MQTT brokers such as EMQX and VerneMQ expose received-versus-delivered counters, which measure loss inside the broker and pipeline but not loss between device and broker. The third is expected-rate modelling: a registry of devices with known reporting intervals gives an expected count per window, which is the weakest option but works for legacy fleets.
Define “valid” carefully. Messages from devices in a declared maintenance state, from decommissioned devices, and malformed messages rejected by schema validation are not platform failures, though the last category deserves its own counter and a separate dashboard.
SLI 2: Ingestion latency
Ingestion latency is the time from the message being received at the edge of your platform (broker publish acknowledgement) to being queryable in the time-series store. It is a classic threshold SLI: a measurement is good if it falls under a bound, for example 5 seconds. A histogram metric on the writer, with a bucket aligned to your threshold, gives you the good-event counter directly, and Prometheus-style le buckets make this cheap. Choose the threshold from what downstream consumers need, not from what the pipeline currently achieves.
Use event timestamps carefully. Devices have unreliable clocks, so measure with the platform’s own receive timestamp at the front and the commit timestamp at the back. Device-supplied timestamps belong in the freshness SLI, where clock skew is part of what you want to detect.
SLI 3: Data freshness
Freshness answers “how old is the newest value the twin is showing?” It differs from latency because it is evaluated continuously, not per message. A pipeline can have excellent per-message latency and still serve stale data if a partition stalls and nothing is being processed at all. Latency SLIs only observe messages that arrived; freshness observes silence.
The standard construction is a time-slice SLI. Every 30 seconds, for each monitored entity group, evaluate whether the age of its newest point is under the freshness bound. A slice is good if age is under, say, two reporting intervals. The SLI is good slices over total slices. OpenSLO supports this directly through its time-slice budgeting method, which we will use below. Aggregate by group (a plant, a region, a tenant tier) rather than per device, so that a few offline devices do not drown out a platform outage.
SLI 4: Device-to-dashboard latency
This is the end-to-end number: the time from a reading being published by a device to its appearing in a query result or a rendered twin. It is the one stakeholders understand, and the one hardest to instrument, because it crosses ownership boundaries. The practical technique is synthetic canaries: a handful of canary devices (real or simulated) publish a heartbeat with a known timestamp every 10 to 30 seconds, and a probe queries the dashboard API and records the delay until each heartbeat becomes visible. Because the canaries run continuously, this SLI has dense traffic even when the real fleet is quiet, which becomes important later in the low-traffic discussion.
Canaries do not replace the real-traffic SLIs. They prove the path works, not that every tenant’s data flows through it. Treat them as the end-to-end tripwire and keep the per-hop SLIs as the diagnosis layer.
Choosing targets without inventing them
Targets should be negotiated, not copied. A useful method is to compute what the last 90 days of data would have scored, set the first objective slightly below the observed level, and tighten only when the business asks for it. A 99.99 percent target is rarely justified for telemetry that is already sampled at 10-second intervals, because the next reading repairs most single-message losses. The arithmetic of what each nine costs is worth seeing once. Over a 30-day window of 43,200 minutes, 99.9 percent allows 43.2 minutes of full outage, 99.5 percent allows 216 minutes (3.6 hours), and 99 percent allows 432 minutes (7.2 hours). Choose the nine that matches the consumer’s tolerance, and record why.
Burn-Rate Math and Multiwindow Alerting in Depth
Burn rate is the speed at which you consume error budget, normalized so that 1.0 means you will exhaust the budget exactly at the end of the SLO window. If the SLO is 99.9 percent, the error budget is 0.001. A burn rate of 1 means the observed error ratio equals 0.001. A burn rate of 14.4 means the observed error ratio is 0.0144, or 1.44 percent. The formula is simply the observed error ratio divided by the budget fraction, where the budget fraction is one minus the objective.
The useful derived quantity is the fraction of the total budget consumed during an alert window. Over a 30-day window (720 hours), a burn rate B sustained for W hours consumes B times W divided by 720 of the budget. For B equal to 14.4 and W equal to 1 hour, that is 14.4 divided by 720, which is 0.02, or 2 percent. For B equal to 6 and W equal to 6 hours, that is 36 divided by 720, which is 5 percent. For B equal to 1 and W equal to 72 hours (3 days), that is 72 divided by 720, which is 10 percent. These three pairs are exactly the table in the Google SRE Workbook for a 99.9 percent SLO, and they are the reason the numbers 14.4, 6 and 1 appear in nearly every burn-rate tutorial.
Why two windows per alert
A single long window fires correctly but resets slowly. If you alert on a one-hour error ratio above 1.44 percent, an incident that lasts five minutes at a very high error rate can keep the alert firing for most of the following hour after it is fixed. The Workbook fixes this by requiring a short window, one-twelfth of the long window, to also exceed the threshold. The alert fires only when the long window says “this is significant” and the short window says “this is still happening”. Per the Workbook, the alert stops firing about five minutes after the problem is resolved rather than an hour later.

Figure 3: A burn-rate alert fires only when both the long-window and short-window error ratios exceed burn rate multiplied by the budget fraction.
The decision in Figure 3 maps to a PromQL expression of the form: error ratio over 1h greater than 14.4 times the budget, AND error ratio over 5m greater than 14.4 times the budget. The three severity tiers cover three different failure shapes.
| Severity | Long window | Short window | Burn rate | Budget consumed if sustained | Typical shape |
|---|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4 | 2 percent | Outage or severe regression |
| Page | 6 hours | 30 minutes | 6 | 5 percent | Significant partial degradation |
| Ticket | 3 days | 6 hours | 1 | 10 percent | Slow leak, chronic low-grade loss |
The table is the Workbook’s recommendation for a 99.9 percent SLO over 30 days. The burn rate thresholds are independent of the objective: at 99.5 percent the budget fraction is 0.005, so the 14.4x page fires at a 7.2 percent error ratio, whereas at 99.99 percent it fires at 0.144 percent. This is the property that makes burn-rate alerting portable. One alert design works for every SLO because it is expressed in multiples of the budget.
Time to detection, worked through
Consider an illustrative IoT case: an ingest pipeline drops 100 percent of messages for one tenant tier that represents all traffic in the scoped SLI, against a 99.9 percent objective. The error ratio is 1.0, which is a burn rate of 1,000. The 14.4x page condition on the 5-minute window is satisfied almost immediately, and the 1-hour window crosses 1.44 percent after about 52 seconds of total outage (1.44 percent of 3,600 seconds). Detection is therefore under about a minute plus the evaluation interval, and the page fires quickly.
Now a partial failure: 2 percent message loss from a firmware bug in one gateway model. That is a burn rate of 20 against a 99.9 percent objective. The 14.4x page fires once the 1-hour ratio crosses 1.44 percent, which with a steady 2 percent rate takes roughly 43 minutes of the hour window filling (1.44 divided by 2 times 60 minutes). The budget consumed by then is about 2 percent. If instead the loss is 0.8 percent (burn rate 8), the 14.4x page never fires but the 6x page does after the 6-hour window crosses 0.6 percent, which takes about 4.5 hours, consuming roughly 5 percent of the monthly budget. A loss of 0.15 percent (burn rate 1.5) is caught only by the ticket tier, after a day or two. Each tier trades detection speed for sensitivity, and the numbers above are arithmetic, not measurements.
Why not just threshold the error rate?
A static threshold on error rate has to be tuned per service and ignores the budget. Thresholds low enough to catch slow leaks page constantly on transient blips, and thresholds high enough to be quiet miss the leaks. Burn-rate alerts are precise because the long window averages out noise and the short window confirms recency, and the severity tiers match human response: wake someone for fast burns, open a ticket for slow ones. The Workbook describes the progression from naive alerts to the final design, showing the precision and recall trade at each step, and it is worth reading in full.
OpenSLO and Sloth: From Specification to Prometheus Rules
The second half of the pattern is operational. You want SLOs defined as code, reviewed in pull requests, validated in CI, and compiled into monitoring rules without hand-writing PromQL. OpenSLO and Sloth split that work cleanly: OpenSLO is a vendor-neutral description of the objective, and Sloth is a compiler that targets Prometheus.

Figure 2: The SLO-as-code pipeline. A spec goes through Sloth generate and becomes three groups of Prometheus rules feeding Alertmanager and Grafana.
An OpenSLO document for data freshness
OpenSLO’s current API version, per the project README, is openslo/v1. An SLO declares a service, a budgeting method (Occurrences, Timeslices or RatioTimeslices), one time window, an indicator and one or more objectives. The README’s own example uses a Prometheus metric source with a ratio metric. Below is an adapted document for the IoT freshness SLI as a time-slice. The metric names (twin_entity_age_seconds and so on) are illustrative, not from any specific product, and you would substitute your own exporter.
apiVersion: openslo/v1
kind: SLO
metadata:
name: twin-data-freshness
displayName: Twin data freshness
spec:
service: twin-state-service
description: Newest point per plant group is younger than 60 seconds
budgetingMethod: Timeslices
timeWindow:
- duration: 30d
isRolling: true
indicator:
metadata:
name: freshness-under-60s
spec:
thresholdMetric:
metricSource:
type: Prometheus
spec:
query: max(twin_entity_age_seconds{tier="production"})
objectives:
- displayName: Fresh data
op: lte
value: 60
target: 0.995
timeSliceTarget: 0.995
timeSliceWindow: 1m
The structure reads as: every one-minute slice, evaluate the query; the slice is good when the value is at most 60; the objective is that 99.5 percent of slices are good over a rolling 30 days. I have composed this example from the documented field names (budgetingMethod, timeWindow, thresholdMetric, timeSliceWindow). Field details for time-slice objectives have changed between OpenSLO versions, so run the document through the OpenSLO tooling or Sloth’s validation before relying on it. I did not execute this specific file in this research pass.
For comparison, a message-loss SLI is a ratio over counters and maps to the Occurrences method, the same shape as the README example. Good is sum(increase(ingest_points_written_total[...])), total is sum(increase(ingest_points_expected_total[...])), where the expected series comes from one of the denominator strategies discussed earlier. Using the Occurrences method means every message counts equally, which is correct for loss, whereas Timeslices treat every minute equally, which is correct for freshness.
The Sloth spec and what it generates
Sloth’s native format is more compact than OpenSLO and is what its documentation leads with. The documented getting-started example looks like this, shown here adapted to an ingest pipeline. The {{.window}} placeholder is Sloth’s template variable, replaced with each window it needs.
version: "prometheus/v1"
service: "ingest-pipeline"
labels:
owner: "platform-team"
tier: "1"
slos:
- name: "message-delivery"
objective: 99.9
description: "Messages accepted by the broker that are durably written."
sli:
events:
error_query: sum(rate(ingest_points_dropped_total{pipeline="main"}[{{.window}}]))
total_query: sum(rate(ingest_points_received_total{pipeline="main"}[{{.window}}]))
alerting:
name: IngestMessageLossBurn
labels:
category: "data-quality"
annotations:
summary: "Ingest message loss is burning error budget"
page_alert:
labels:
severity: page
routing_key: platform
ticket_alert:
labels:
severity: ticket
Running sloth generate -i spec.yml -o rules.yml emits standard Prometheus rule groups. According to the Sloth documentation example, the output has three kinds of groups per SLO: SLI recordings (named like sloth-slo-sli-recordings-<service>-<slo>) that compute the error ratio at 5m, 30m, 1h, 2h, 6h, 1d, 3d and 30d windows; metadata recordings (sloth-slo-meta-recordings-...) that publish the objective, the error budget and burn-rate gauges; and alert rules (sloth-slo-alerts-...) with a page alert and a ticket alert using the multiwindow, multi-burn-rate logic.
From memory of Sloth’s generated output, the recording rule metric names follow the pattern slo:sli_error:ratio_rate5m (and ..._rate30m, ..._rate1h and so on), with slo:objective:ratio, slo:error_budget:ratio, slo:current_burn_rate:ratio, slo:period_burn_rate:ratio and slo:period_error_budget_remaining:ratio as metadata series, labelled with sloth_service, sloth_slo and sloth_id. I could not confirm these exact metric names on the pages fetched in this research run, so treat them as unverified and diff them against your own sloth generate output before building dashboards on them.
The generated alert is conceptually the Workbook design in rule form. A trimmed illustration, not verbatim Sloth output, looks like this:
- alert: IngestMessageLossBurn
expr: |
(
slo:sli_error:ratio_rate1h{sloth_id="ingest-pipeline-message-delivery"} > (14.4 * 0.001)
and
slo:sli_error:ratio_rate5m{sloth_id="ingest-pipeline-message-delivery"} > (14.4 * 0.001)
)
or
(
slo:sli_error:ratio_rate6h{sloth_id="ingest-pipeline-message-delivery"} > (6 * 0.001)
and
slo:sli_error:ratio_rate30m{sloth_id="ingest-pipeline-message-delivery"} > (6 * 0.001)
)
labels:
severity: page
Two design points follow from this. Because the heavy computation happens once in recording rules, the alert expressions are cheap and evaluate in constant time regardless of how many devices feed the underlying counters. And because the 30-day window is also recorded, dashboards can show the remaining error budget without running a 30-day range query on every refresh.
Window and period configuration
Sloth ships two period catalogs: a 30-day window aligned with the SRE Workbook and a 28-day alternative built from four whole weeks. A 28-day window has an operational appeal because it always contains exactly four of each weekday, which matters for fleets with weekly production cycles; a 30-day window mixes in five of some weekdays. You can override the default with --default-slo-period or define your own catalog (for example 7 days) through the AlertWindows/v1 spec and the --slo-period-windows-path flag. A 7-day window gives faster budget recovery and tighter feedback but amplifies noise, so reserve it for fast-moving services.
Operating modes
Sloth can run as a CLI in CI, generating rule files committed or deployed by GitOps, or as a Kubernetes controller that reconciles PrometheusServiceLevel custom resources into Prometheus Operator rules. The release listing for v0.16.0 also mentions a server command with a UI for browsing services and SLO statistics, and a validate command that detects duplicate SLOs. For a platform team that already stores everything in Git, the CLI plus CI validation is the lowest-friction start, and the controller becomes attractive once application teams own their own SLOs.
Error Budget Policy: Making the Number Change Behavior
An SLO that no one acts on is a dashboard decoration. The error budget policy is the document that says what happens at each level of budget consumption, and it must be agreed by engineering and product before the first incident, not during it. The Google SRE Workbook includes an example policy template, and the shape is consistent: thresholds, consequences, an escalation path, and an exception process.

Figure 4: An example error budget policy. The thresholds are illustrative and should be negotiated per service.
A workable policy for an IoT platform has four elements. First, a scope: which SLOs it governs and which team owns each. Second, a gradient of consequences. Above roughly half of the budget remaining, ship at normal pace. Below that, require extra review and canary stages for changes touching the ingest path. At zero, freeze feature releases for that service until the budget recovers, except for fixes that improve reliability or security. The 50 percent figure is an example, not a standard. Third, a mandatory postmortem trigger: any single incident that consumes more than a stated share of the monthly budget (many teams choose 20 percent or so) gets a written review with tracked actions. Fourth, an exception route, because a rigid freeze applied to a regulatory fix does more harm than good.
What the policy does in an IoT context
IoT release processes have an extra wrinkle: firmware. A firmware rollout to a million devices is a change with a long tail and no instant rollback, so it deserves its own gating rule in the policy. A reasonable rule is that fleet-wide firmware pushes require a healthy budget on the ingestion and loss SLOs, because a bad build will burn both. Staged rollout rings (1 percent, 10 percent, 50 percent) fit naturally: the SLO dashboard filtered to each ring is the promotion criterion, and a burn-rate page during a ring is the automatic halt signal.
The policy also helps with the platform-versus-device blame problem. If a customer’s gateway has a flaky cellular link, that is not your budget. Define in the SLI which failures count (platform-side drops, pipeline delays, store write errors) and which do not (device offline, customer network). Excluding them is legitimate only if you measure them separately and show them. A hidden exclusion list is how SLOs become dishonest.
Connecting burn-rate pages to response
A page needs a destination. For fast burns, route to the on-call engineer with a runbook that starts from the inner SLIs: which hop is burning? For ticket-tier burns, route to the owning team’s backlog with the dashboards linked. Teams building automated triage on top of these signals can feed burn-rate alerts into an incident workflow; the architecture for that is discussed in AIOps incident response architecture and agentic SRE, where a well-defined SLO page is a far better trigger than a raw metric threshold because it carries business impact and severity by construction.
Deeper Analysis: Making SLIs Work on Real IoT Traffic
The textbook recipe assumes dense, roughly stationary traffic. IoT fleets violate that in several ways, and each needs a deliberate answer.
Low-traffic and bursty denominators
Burn-rate alerts divide errors by total events. When total traffic is tiny, a single failure can look like a 100 percent error ratio over the 5-minute window and trip the page. The Workbook has a dedicated discussion of low-traffic services, with options such as generating artificial traffic, combining services into a larger SLI, and modifying the client to retry. For IoT, the best answer is usually combination plus canaries: aggregate the SLI over a tenant tier or region rather than per device, and keep the synthetic canary stream running so that the end-to-end SLI always has a floor of events.
The opposite problem is the reconnect storm. After a broker restart, every device reconnects and republishes buffered data within minutes. Total volume spikes, the ratio dilutes, and a real problem can hide in the surge. Weighting by event count is still correct, but you should watch the absolute error count alongside the ratio during recovery, and make sure your pipeline sheds load predictably (explicit backpressure and retry-after semantics) rather than dropping silently.
Late-arriving data and event time
Devices on intermittent links deliver data late, sometimes hours late. If your freshness or loss SLI is evaluated in processing time, late data looks like loss first and then like recovery. Decide which clock the SLI uses and document it. For freshness, use processing time: the question is “how old is the data we are showing now”. For completeness, use event time with a grace period: a message that arrives within, for example, the configured late tolerance counts as delivered. Anything beyond that is either loss or an accepted exclusion, depending on policy.
Cardinality and cost
Per-device SLIs are tempting and expensive. A million devices with a label each means a million series per metric, multiplied by histogram buckets. The pattern that scales is to compute SLIs from aggregates: counters summed per tenant tier and pipeline stage in the ingest layer, and slices evaluated per entity group. Sloth’s recording rules depend on the aggregated queries you give it, so write the sum in the query, not downstream. Keep device identifiers out of SLI labels; keep them in logs and traces where they belong, and use the observability patterns from the OpenTelemetry post linked earlier to join them when diagnosing.
A rough capacity illustration makes the point. Suppose a fleet of 50,000 devices each reports every 10 seconds. That is 5,000 messages per second, or 12.96 billion messages in a 30-day window (5,000 times 2,592,000 seconds). A 99.9 percent loss objective then allows 12.96 million lost messages per month. Those numbers are arithmetic on assumed inputs, not a benchmark, but they show why the budget is best tracked as a ratio recorded in Prometheus rather than recomputed from raw events: the raw event volume is too large to scan on every dashboard load.
Per-tenant fairness and skew
An aggregate SLI can look healthy while one customer suffers. If a tenant representing 2 percent of traffic loses everything, the fleet-wide error ratio is 2 percent, which is a burn rate of 20 against a 99.9 percent objective and does fire the page, but at 0.5 percent of traffic it is a burn rate of 5 and falls into the 6-hour tier or nothing. For contractual tenants, define SLOs on tenant tiers, or run the same SLO template per tenant for the top accounts. Sloth’s labels and per-SLO generation make this a template exercise rather than a manual one, but the cost grows linearly with the number of SLOs, so tier the customers.
Choosing between Occurrences and time slices
The budgeting method changes what an outage costs. With Occurrences, a ten-minute outage during the busiest period burns more budget than during a quiet period, because more events fail. With Timeslices, every minute is equal. Message loss belongs in Occurrences because each message matters. Freshness belongs in time slices because it is a state of the system, not a count of events. Mixing them up is a common mistake, and it produces budgets that do not match intuition about the harm done.
Trade-offs, Gotchas, and What Goes Wrong
Too many SLOs. Teams that discover SLOs often create dozens per service. Each needs a target, an owner and a response path, and alert fatigue returns through the back door. Start with one to three SLOs per user-facing journey: for a twin platform, delivery, freshness and end-to-end latency are enough.
Measuring at the wrong place. An SLI computed from the server’s own logs misses everything lost before the server. Measure as close to the consumer as you can, and where you cannot, say what you are not covering. A freshness SLI built on the twin service’s internal timestamps will happily report good health while the dashboard API is down.
Targets copied from a book. A 99.99 percent objective means 4.32 minutes of budget per 30 days, which is below the resolution of most deployment pipelines and human response. If the budget cannot be spent, the policy cannot work. Set the objective from the consumer’s tolerance and your observed baseline.
The 14.4x number applied blindly. The 14.4, 6 and 1 multipliers are derived for a 30-day window and the specific budget percentages in the Workbook (2, 5 and 10 percent). If you change the window to 7 days, the factors that consume the same budget percentages change too: for a 7-day window of 168 hours, a 2 percent budget burn in one hour is 0.02 times 168, or 3.36, not 14.4. Sloth’s window catalog handles this when you configure a custom period; hand-written rules often do not.
Silent exclusions and gaming. If an SLI excludes whatever is currently broken, it will always look green. Review the exclusion list in the same pull request process as the SLO itself, and publish the excluded events as their own metric.
Alerting on an SLO with no history. Brand-new SLOs have no 30-day window, so the long-window recording rules are empty or partial. Sloth’s alert rules are built to cope, but dashboards showing remaining budget will be misleading for the first weeks. Annotate them as warming up.
Dependence on Prometheus semantics. Counter resets, scrape gaps and staleness markers all affect rate(). If an ingest pod restarts and its counter resets, rate handles it, but a scrape gap during an incident can understate errors. Pair the SLO with a meta-alert on the monitoring pipeline itself: absence of the SLI series should alert.
OpenSLO compatibility is not magic. OpenSLO is a specification, and not every tool supports every field. Sloth supports OpenSLO as one of its input formats, but check which budgeting methods and indicator types your generator accepts before standardizing on the spec. Some teams author in Sloth’s native format for exactly this reason and keep OpenSLO as an export target.
Practical Recommendations
Begin with the data path, not the tool. Draw the hops from device to dashboard, decide the consumer-visible failure at each, and pick the three or four SLIs that cover them. Run the candidate SLIs in a dashboard for four weeks before attaching alerts, so you learn the real baseline and discover denominator problems while the stakes are low.
Then codify. Write the SLOs in OpenSLO or Sloth’s native spec, keep them in the same repository as the platform’s infrastructure code, and validate in CI with sloth validate. Generate rules in the pipeline and deploy them with the rest of your monitoring configuration. Keep page routing conservative: only the two fast tiers page a human, and the 3-day tier opens a ticket.
Finally, make the policy real. Hold a short review meeting every month where the budget numbers are read aloud, and record each time the policy changed a decision. If it never changes a decision, the target is too loose or the policy too soft.
A rollout checklist:
- Map the data path and name one owner per hop.
- Define good and valid events for each SLI, including documented exclusions.
- Establish the denominator for message loss (sequence numbers, broker counters or expected-rate model).
- Stand up canary devices for the end-to-end SLI.
- Observe four weeks of baseline before alerting.
- Write SLOs as code, validate in CI, generate rules with Sloth.
- Page on 14.4x (1h and 5m) and 6x (6h and 30m) burns; ticket on 1x (3d and 6h).
- Sign the error budget policy, including the firmware rollout gate.
- Alert on missing SLI series and review exclusions monthly.
Frequently Asked Questions
What is an SLO error budget?
An error budget is the amount of unreliability an SLO permits, equal to one minus the objective. For a 99.9 percent SLO over 30 days, the budget is 0.1 percent of valid events, or about 43.2 minutes of full downtime in time-based terms. Teams spend it on releases, experiments and incidents. When it is exhausted, policy slows or stops risky changes until reliability recovers. It turns reliability into a shared, measurable resource instead of a debate.
What is burn rate and why is 14.4 the standard page threshold?
Burn rate is the observed error ratio divided by the budget fraction, so a value of 1 exhausts the budget exactly at the end of the window. On a 30-day window, a 14.4x burn sustained for one hour consumes 2 percent of the budget (14.4 times 1 hour divided by 720 hours). The Google SRE Workbook chose 2 percent in one hour as the level that warrants an immediate page, which yields 14.4.
How do I measure data freshness for an IoT platform?
Treat freshness as a time-slice SLI. At a fixed interval, such as 30 or 60 seconds, check whether the age of the newest data point for each entity group is under a bound, typically two reporting intervals. A slice is good if the age is under the bound. The SLI is good slices divided by total slices. Aggregate by plant, region or tenant tier rather than per device, to avoid noise from individual offline devices.
Should I use OpenSLO or Sloth?
They solve different problems and work together. OpenSLO is a vendor-neutral specification for describing SLOs, useful when you want portability across tools. Sloth is a generator that compiles SLO definitions into Prometheus recording and alerting rules implementing multiwindow burn-rate alerts, and it can read OpenSLO input as well as its own format. If you run Prometheus, using Sloth gives you working alerts quickly, and OpenSLO keeps your definitions portable.
How do I handle SLOs when traffic is very low?
Low traffic makes ratios volatile, because one failed event can look like a total outage. Options include aggregating across tenants or regions into one larger SLI, running synthetic canary traffic so the SLI always has events, using longer windows for noisy services, or accepting a ticket instead of a page. Avoid per-device SLOs for paging purposes, and monitor absolute error counts alongside ratios during quiet periods.
Do I need a 30-day window?
No, but it is the default in the Google Workbook and Sloth, and the standard burn-rate multipliers are derived for it. Sloth also ships a 28-day catalog, which always contains four of each weekday, and lets you define custom windows such as 7 days. If you change the window, the burn-rate factors and budget percentages must be recomputed, which Sloth does for its catalog entries.
Further Reading
- AIOps incident response architecture and agentic SRE for routing burn-rate pages into automated triage.
- OpenTelemetry versus Prometheus and Loki for edge observability for deciding how the underlying telemetry is collected.
- OTLP versus Prometheus remote write for the transport choice behind SLI metrics.
- Google SRE Workbook: Alerting on SLOs, the source of the multiwindow burn-rate design.
- Sloth documentation and the OpenSLO specification.
By Riju — about
