Usage-Based Billing Architecture: Metering, Rating, Invoicing and Idempotent Events
Educational systems analysis only. Nothing here is financial, tax, legal or accounting advice.
A flat subscription fails quietly: the price is a constant, so the billing code is a cron job. A usage-based price fails loudly. Every API call, gigabyte or token becomes a financial fact, and one dropped, doubled or misplaced event is a wrong invoice that a customer can read line by line. Teams that move from seats to consumption discover that billing stops being a feature and becomes a distributed data system with an audit requirement.
This is why usage-based billing architecture deserves the same rigour as a payments ledger. The hard parts are not the arithmetic. They are accepting events at-least-once and still billing exactly once, deciding what to do with data that arrives after the books closed, and proving afterwards that the invoice equals the evidence.
You will leave with a five-stage reference pipeline, an event schema with idempotency keys, a watermark policy for late data, a tested graduated-pricing function, an SQL aggregation, an invoice-correction model and a build-versus-buy matrix.
What this covers: the metering-to-collection pipeline, idempotent ingestion, late and out-of-order events, aggregation windows, tiered pricing math, proration, credits and commitments, invoice immutability, reconciliation, failure modes and build-versus-buy.
Context and Background
Consumption pricing is old. Telephone and electricity utilities have run call detail records and meter reads through mediation, rating and billing for decades; the telecom industry even standardised the vocabulary that modern SaaS billing quietly reuses. What changed is the producer population. A utility had a few million meters reporting on a schedule. A cloud product has thousands of services emitting millions of events per second, from code that deploys several times a day and retries on any timeout.
That shift moved the problem from batch file processing to streaming ingestion, and it moved the failure mode. In telecom a malformed record file was rejected whole. In a modern service mesh a single misconfigured emitter can double every count for an afternoon while every health check stays green. The billing system is often the only place where such a bug becomes visible, and by then it is a customer dispute.
Today there are three broad ways to cover the problem. You can use the metering features of a payment platform; Stripe Billing, for instance, accepts meter events through an API and aggregates them against a configured meter, and its documentation on recording usage is a good statement of the practical limits of that approach. You can adopt an open-source engine; Lago is described by its maintainers as an open-source billing platform under the AGPLv3 licence, and OpenMeter is Apache 2.0 licensed and ingests CloudEvents into a Kafka and ClickHouse based runtime. Or you can build the pipeline yourself, which is what the rest of this article is a blueprint for, even if you eventually buy.
Idempotency is the idea everything hangs on. If you have not designed retry-safe write paths before, our guide to idempotent API design covers the mechanics of keys and replay windows that this article applies to events. Likewise, the final control in the pipeline is the same discipline used for payments, covered in reconciliation engine architecture.
A final piece of context is cultural. Finance, product and engineering use the same words differently. “Usage” to a product manager is a dashboard number; to finance it is revenue that must tie to a general ledger; to a customer it is a promise on an invoice. A good architecture makes the three views derive from one immutable record rather than from three pipelines that drift apart.
The Reference Pipeline: Metering, Aggregation, Rating, Invoicing, Collection
The short answer: a usage-based billing system is a pipeline of five stages that each narrow ambiguity. Metering captures immutable usage events. Aggregation folds them into per-customer, per-window quantities. Rating converts quantities into money using a versioned price book. Invoicing freezes that money into a legal document. Collection moves the cash and reconciliation proves the books match the provider.

Figure 1: The five-stage pipeline. The raw event store is the single source of truth; every downstream number can be rebuilt from it.
The diagram shows a deliberate asymmetry. Everything after the durable log is derived and therefore rebuildable, while the log itself is append-only. If the rating logic has a bug, you fix the code and replay. If the log is wrong, you have lost evidence. That one design rule drives most decisions below.
Stage one: metering captures facts, not prices
A metering event states that something happened: customer C consumed N units of metric M at time T. It carries no price, no tier and no currency. Keeping money out of events is what lets you change a price book, correct a rating bug or run a what-if simulation without touching history. Events should also be emitted as close as possible to the work that was actually performed, because a gateway that counts requests will bill a customer for requests that later failed.
Two timestamps matter. Event time is when the usage occurred; ingestion time is when your system learned of it. Billing windows are defined on event time, while operational concerns such as lag and watermarks are defined against ingestion time. Confusing them is the root cause of most disputes about late data.
Stage two: aggregation folds events into quantities
Aggregation maps many events to one number per customer, metric and window. The aggregation function is part of the product definition, not an implementation detail. Sum suits bytes transferred or tokens consumed. Count suits API calls. Unique count suits monthly active users. Maximum or last-value suits provisioned capacity such as the peak number of seats or the final storage level. Stripe’s documentation, for example, lets a meter use different aggregations, and its API accepts decimal values, which is useful for fractional units such as gigabyte-hours.
The window is equally explicit: calendar month, the subscription anniversary period, or a rolling window for entitlements. Aggregation must be deterministic so that replaying the same events returns the same quantities.
Stage three: rating applies a versioned price book
Rating is a pure function from quantity, price plan and context to money. Pure matters: given the same inputs it must return the same output, which means no reading the current time, no hidden mutable discount tables, and no floating-point money. A price book is therefore versioned and effective-dated, and each invoice records which version was used. This lets finance answer, months later, why a customer paid what they paid.
Stage four: invoicing freezes the result
An invoice is a document with legal weight in most jurisdictions. It has a number, a period, line items, taxes and a total, and once issued it should not change. The pipeline treats the invoice as the point where derived numbers become records. Anything discovered afterwards is handled by additional documents (credit notes and corrected invoices), not by editing.
Stage five: collection and reconciliation
Collection hands the invoice to a payment provider and tracks the outcome: paid, failed, disputed or refunded. Reconciliation then compares your ledger with the provider’s settlement data. Without it you have an elegant pipeline that nobody can prove correct.
Why log first, derive later
A tempting shortcut is to increment a counter in a database row per customer for each event. It is fast and simple, and it fails exactly when you need it. A counter cannot tell you which events it contains, so you cannot dedupe, cannot replay after a rating bug, and cannot explain a number to a customer. An append-only log of events plus derived aggregates costs more storage and gives you provenance. In practice the storage cost is modest compared with a single credit-note storm.
There is a second reason to keep the log: regulators and auditors ask to see evidence. Retention periods for billing records are set by local tax and commercial law and vary widely, so the right answer is to ask your finance and legal teams, then keep raw events at least that long in cheap object storage with the aggregates in a faster store.
Event Schema and Idempotent Ingestion
The short answer: every usage event needs a producer-generated unique identifier, a customer, a metric, a quantity and an event time. The ingestion layer deduplicates on the identifier within a retention window, so at-least-once delivery from clients still yields exactly-once effect in the totals. Idempotent ingestion is a property of the pair of producer and store, not of either alone.
A minimal event schema
The CNCF CloudEvents specification is a useful starting point because it already encodes the identity rule. It requires id, source, specversion and type, and states that producers must ensure source plus id is unique for each distinct event, while consumers may treat events with identical source and id as duplicates. Note the word “may”: the specification does not enforce deduplication, so that is your job. OpenMeter, for example, accepts CloudEvents with a subject attribute naming the customer and a data payload holding the measured value.
A pragmatic billing event looks like this:
{
"specversion": "1.0",
"id": "9f1c2a7e-5d0b-4c53-9a2e-3b6f1e0a77d1",
"source": "svc/inference-gateway",
"type": "tokens.consumed",
"subject": "cust_8841",
"time": "2026-10-09T14:03:22.418Z",
"data": { "value": 1840, "unit": "token", "model": "m-large", "region": "ap-south-1" }
}
Several choices here are deliberate. The identifier is generated by the producer at the moment of the business action, not at send time, so a retry carries the same value. The customer is a stable internal key, never an email address. The data.value is numeric and carries a unit so the rating layer can reject a gigabyte reported where a megabyte was expected. Dimensions such as model and region are the basis for per-dimension pricing, and they need governance: Stripe documents limits on unique dimension combinations per meter and per customer, which is a hint that unbounded dimension values (a user identifier used as a dimension, say) turn a metering system into a cardinality problem.
The producer side of the contract
Producers must derive the identifier from the business action, not from the transmission attempt. The safest patterns are a UUID minted when the work is recorded in a local outbox table, or a deterministic hash of stable inputs such as request identifier plus metric plus sequence number. A UUID minted inside the HTTP client library on each attempt is the classic bug: every retry looks like a new event and every timeout inflates the bill.
The producer should write the event to a local transactional outbox in the same database transaction as the work, then ship it asynchronously. This closes the gap where work succeeds but the emit call fails, which would otherwise create unbilled usage. The same pattern appears in any system that must update state and publish a message atomically.
The ingestion side of the contract

Figure 2: A retried event is accepted, recognised by its identifier and ignored, so totals stay correct even though the producer saw a timeout.
The ingest API validates the schema, checks the timestamp against sanity bounds, performs an insert-if-absent on the identifier and appends to the log. Responding with the same success code for a duplicate as for the original matters: a client that gets an error for a duplicate will retry again and never converge. Stripe’s meter events behave in this spirit: every event has an identifier that you can set, and if you do not, one is generated, which means a retry without a stable identifier is not deduplicated.
Stripe also enforces timestamp bounds that illustrate sane defaults: an event timestamp must be within the past 35 calendar days and no more than 5 minutes in the future, the latter to tolerate clock drift. Even if you build your own, define both bounds. Future timestamps are usually clock bugs, and unbounded past timestamps let a backfill silently rewrite a closed month.
Where to keep the dedup state
Dedup state is a set of identifiers with a time-to-live. Three designs are common. A relational unique index on (tenant, source, id) is simplest and gives strong guarantees at moderate volume. A key-value store with set-if-absent and TTL scales further but must be sized so that eviction never occurs inside the retry horizon. A log-compaction approach lets the log itself deduplicate by key, trading latency for operational simplicity. OpenMeter’s repository mentions optional Redis for distributed deduplication; the details of its algorithm are not in the pages I could verify, so check its architecture guide before relying on specifics.
The dedup horizon is a business decision with a numeric consequence: if producers can retry for 24 hours, your window must exceed 24 hours with margin, and an event that arrives after the window can be duplicated. Document that residual risk and cover it with reconciliation on the aggregate.
The exactly-once myth
Marketing copy talks about exactly-once delivery. Networks give you at-most-once or at-least-once, and what systems call exactly-once is at-least-once delivery plus idempotent processing, producing an exactly-once effect. Message brokers with transactional features narrow the window inside the broker, but they do not cover the hop from your producer through your API or the hop into your database. The honest claim is: duplicates are tolerated, and the effect on totals is counted once. Treat any component that promises more with suspicion, and test it by killing processes mid-flight.
A practical test is a chaos replay: send a recorded day of events through the system with random connection resets and double-submits, then assert that the aggregates equal those from a clean single pass. If they differ, you have found a bug before a customer did.
Authentication and tenancy
Usage events are money. An emitter must be authenticated and authorised for the customers it can report, otherwise one compromised service can inflate or erase another tenant’s bill. Scope credentials per source, include the source in the dedup key so two services cannot collide on identifiers, and rate-limit per source. Stripe applies a concurrency limit of one concurrent call per customer per meter, which shows another design option: serialising writes per customer to avoid races in the aggregate.
Late, Out-of-Order and Corrected Events
Real streams are messy. A mobile client buffers events for hours offline. A cluster partitions and flushes backlog. A batch job reports yesterday’s usage today. Out-of-order arrival is normal, so the system needs a policy, not an exception handler.
Event time, ingestion time and watermarks
A watermark is the system’s claim that it has seen all events with event time earlier than some point. It is a heuristic: you choose it based on how late data typically is, and you accept that some events will be later still. A conservative watermark delays invoice finalisation; an aggressive one produces more late records. The right setting is the 99.9th percentile of observed lag plus a safety margin, measured, not guessed. Because lag distributions are heavy-tailed, report both the median and the tail on a dashboard and alert when the tail grows.
Windows close only when the watermark passes the window end plus a grace period. During the grace period the draft invoice for that window remains open and can be re-rated as late events arrive. Stripe offers a similar concept as an invoice finalisation grace period, and its documentation notes that an event can be accepted yet not included on an already-finalised invoice, which is the key insight: acceptance and billability are separate states.

Figure 3: Late-event policy. The decisive question is whether the invoice for the event’s window has already been finalised.
The three outcomes for a late event
When an event’s time falls before the watermark there are exactly three coherent choices.
First, absorb: the window is still open, so the event simply joins the aggregate and the draft is re-rated. This is free and invisible, and it is why a grace period is worth its delay.
Second, carry forward: the invoice is finalised, so the event is recorded as a late-usage item and billed on the next invoice with a clear description and its original event time. This is simple and keeps invoices immutable, at the price of a period mismatch that finance must accept.
Third, adjust: the invoice is finalised and the amount is material, so you issue a corrected invoice or a supplementary one. This is the most transparent and the most expensive operationally, and it is typically reserved for amounts above a threshold or for customers whose contracts require period-accurate billing.
Choose per customer class, publish the policy in the order form or terms, and encode it as configuration, not as an engineer’s judgement during close.
Corrections and negative usage
Sometimes the right action is to retract usage: a bug double-counted, a customer was credited after an outage, or an event was emitted for work that was rolled back. The clean way is a compensating event with a negative value, linked to the original identifier, rather than deleting rows. This keeps the log append-only and the audit trail intact. Stripe’s documentation states that if overall cycle usage is negative, the invoice line shows a quantity of zero, which is a useful reminder that your rating layer should define floor behaviour explicitly instead of discovering it from a customer.
Aggregation in SQL
For moderate volumes a relational store can run the aggregation directly. The following PostgreSQL sketch deduplicates by identifier and aggregates a month of usage per customer and metric. It uses a unique constraint for ingestion and a half-open interval on event time, which avoids double-counting an event stamped exactly on a boundary.
CREATE TABLE usage_event (
tenant_id text NOT NULL,
source text NOT NULL,
event_id text NOT NULL,
customer_id text NOT NULL,
metric text NOT NULL,
quantity numeric(20,6) NOT NULL,
event_time timestamptz NOT NULL,
ingested_at timestamptz NOT NULL DEFAULT now(),
dims jsonb NOT NULL DEFAULT '{}',
PRIMARY KEY (tenant_id, source, event_id)
);
CREATE INDEX usage_event_window_idx
ON usage_event (customer_id, metric, event_time);
-- Ingestion is idempotent by construction:
-- INSERT ... ON CONFLICT (tenant_id, source, event_id) DO NOTHING;
-- Monthly aggregate, half-open window [start, end)
SELECT customer_id,
metric,
sum(quantity) AS units,
count(*) AS event_count,
max(ingested_at) AS last_ingested
FROM usage_event
WHERE tenant_id = 'acme'
AND event_time >= timestamptz '2026-09-01 00:00:00+00'
AND event_time < timestamptz '2026-10-01 00:00:00+00'
GROUP BY customer_id, metric;
Two details deserve emphasis. The primary key makes duplicate inserts harmless, so the application can retry blindly. And last_ingested is stored with the aggregate so that, at finalisation, you can record “this invoice reflects all events ingested up to this instant”, which turns a vague claim into a testable one. At higher volumes the same logic moves to a columnar store with pre-aggregated hourly rollups, but the semantics should not change, and a rollup must remain a pure function of the log.
Aggregation windows and time zones
Windows look trivial and cause surprising defects. A calendar month in UTC is not a calendar month for a customer billed in a local zone, and daylight-saving changes make some days 23 or 25 hours long. Anniversary billing, where a period runs from the 15th to the 14th, adds month-length ambiguity: what is the next period after January 31? Decide on a documented rule, store window boundaries explicitly as timestamps on the invoice, and never recompute them from a formula at read time.
Rollups introduce a related trap. If you aggregate hourly and then sum the hours, a late event must update the correct hour bucket, which means rollups need to be recomputable per bucket, not just incremented. Increment-only rollups are fast and drift under late data; recompute-per-bucket rollups cost slightly more and stay correct.
Rating Engine: Pricing Math, Proration, Credits and Commitments
The short answer: a rating engine is a pure function that converts an aggregated quantity into money using a versioned price book. The three tier models differ in how one quantity maps to price: graduated prices each slice at its own rate, volume prices the whole quantity at the rate of the tier it lands in, and flat or per-unit models ignore tiers entirely. Use decimal arithmetic and explicit rounding.
Graduated versus volume pricing, with a worked example
All figures below are illustrative, not any vendor’s real prices. Suppose a price book charges 0.10 per unit for the first 1,000 units, 0.08 per unit for units 1,001 to 10,000, and 0.05 per unit beyond 10,000. A customer consumes 12,500 units in the month.
Under graduated pricing each slice is billed at its tier rate: 1,000 x 0.10 = 100.00; 9,000 x 0.08 = 720.00; 2,500 x 0.05 = 125.00; total 945.00. Under volume pricing the whole 12,500 units fall in the top tier and cost 12,500 x 0.05 = 625.00. The same usage differs by 320.00 depending on the model, which is why the model is a contractual term and not a code detail.
Volume pricing has a known cliff. With these tiers, 10,000 units cost 800.00, while 10,001 units cost about 500.05, so a customer pays less for using more. That is sometimes intended as an incentive, but it invites gaming, and finance should sign off on it. Graduated pricing is monotonic: more usage never costs less.
A runnable, decimal-safe rating function
The following Python implements both models plus proration. It uses Decimal and banker’s rounding, takes no clock or global state, and I ran it to confirm the worked numbers above: graduated(12500) returns 945.00, volume(12500) returns 625.00, graduated(800) returns 80.00 and prorate of a 99 fee over 11 of 30 days returns 36.30.
from decimal import Decimal as D, ROUND_HALF_EVEN
# (upper_bound_inclusive, unit_price); None means unbounded. Illustrative.
TIERS = [(1000, D("0.10")), (10000, D("0.08")), (None, D("0.05"))]
CENT = D("0.01")
def graduated(units, tiers=TIERS):
"""Each slice of usage is priced at its own tier rate."""
total, prev = D(0), 0
for cap, price in tiers:
top = units if cap is None else min(units, cap)
if top > prev:
total += (top - prev) * price
if cap is None or units <= cap:
break
prev = cap
return total.quantize(CENT, ROUND_HALF_EVEN)
def volume(units, tiers=TIERS):
"""The whole quantity is priced at the rate of the tier it lands in."""
for cap, price in tiers:
if cap is None or units <= cap:
return (units * price).quantize(CENT, ROUND_HALF_EVEN)
def prorate(fee, days_used, days_in_period):
return (fee * days_used / days_in_period).quantize(CENT, ROUND_HALF_EVEN)
if __name__ == "__main__":
print(graduated(12500), volume(12500), graduated(800), prorate(D("99"), 11, 30))
Notice what is absent: no database call, no datetime.now(), no float. Purity makes the function trivially testable with property tests: monotonicity of the graduated model, equality of the two models at tier boundaries when prices are equal, and invariance under splitting usage across days when the aggregate is the same. Rounding is the other subtle point. Round once per invoice line, not per event, otherwise a million events at a fraction of a cent each accumulate rounding error. Stripe’s invoices support decimal quantities, which is another reason to keep quantities as numeric and not integers.
Proration
Proration arises when a recurring fee or a commitment covers part of a period: a mid-month upgrade, a signup on the 11th, a cancellation. The common formula is fee x days_used / days_in_period, as in the code above, but there are real choices to document: whether the day of change is counted, whether you prorate by day or by second, and whether a downgrade creates a credit or waits for the next cycle. Usage itself is not prorated, because it is measured, only the fixed component is. A frequent defect is prorating a minimum commitment by calendar days while the usage tiers still reset on the full-period boundary, producing a customer who hits top-tier rates in a partial month.
Credits, commitments and drawdown
Enterprise contracts rarely bill raw usage. A customer may prepay a credit balance that usage draws down, commit to a minimum spend per period, or hold a free allowance that renews monthly. These constructs create ordering questions that the rating engine must answer deterministically.
A reasonable order of operations: compute usage charges from the price book; subtract any included allowance per metric; then apply promotional credits before paid prepaid credits, because promotional credits usually expire sooner; then apply the commitment, billing the shortfall as a true-up line if usage-after-credits is below the minimum; finally compute tax on the amounts the jurisdiction treats as taxable. The order changes the totals, so encode it as an explicit pipeline and test it with golden invoices.

Figure 4: From rated lines to an immutable invoice, with credit drawdown, commitment true-up and a correction path through credit notes.
Drawdown implies a credit ledger: each grant has an amount, an expiry and a priority, and each consumption is an entry linking the invoice to the grant. Treat this as double-entry bookkeeping, so the balance is always the sum of entries and never a mutable field. Expiry is a revenue-recognition event that your finance team will care about, and the rules for breakage vary, so keep that decision with them.
Price book versioning
A price book is a table of plans, metrics, tiers and effective dates. Never update a row in place. Insert a new version with an effective date, and record the version identifier on every rated line. Grandfathering, where existing customers stay on an old version, then becomes a pointer from the customer to a version rather than a code branch. When you need to re-rate history to investigate a dispute, you replay events against the version that was effective, which is only possible if versions are immutable.
Invoice Immutability, Corrections and Reconciliation
What immutable means in practice
Once an invoice is finalised, its lines, amounts, number and period are frozen. The database should enforce it: a status column that moves one way from draft to finalised, with triggers or application guards rejecting updates to finalised rows. Many jurisdictions require sequential invoice numbering without gaps, and many treat issued invoices as legal records, but the specifics differ, so confirm your obligations with your tax adviser rather than assuming.
Immutability is not rigidity. It is what makes corrections auditable. A finalised invoice with an error is corrected by a credit note that references the original, and where required, a new invoice with the right values. The customer and auditor see the original, the reversal and the replacement, each with its own timestamp and reason code. The alternative, silently editing, destroys the trail that finance needs at audit time.
Snapshotting inputs
To make an invoice reproducible, store alongside it the inputs that produced it: the price book version, the aggregate quantities per metric and window, the watermark and last-ingested instant, the credit-ledger entries consumed and the tax rules applied. Then a dispute is a question of recomputing from the snapshot and, if needed, from the raw events, and the two should agree. When they do not, you have found either a late event, a rating bug or a data-quality issue, and the snapshot tells you which.
Reconciliation against the provider
Billing correctness has three layers, and each needs its own check. The first is event integrity: does the count of events accepted by ingestion equal the count produced by emitters? Emitters can publish a periodic control total, such as a per-source count per hour, and the billing system compares it with what it accepted. The second is aggregation integrity: does the sum of the hourly rollups equal the sum of the raw events for the same window? That should be a nightly job that fails loudly. The third is cash integrity: does what you invoiced and collected equal what the payment provider reports?
The third check follows the pattern in our article on reconciliation engine architecture: match ledger entries to provider records by identifier, classify breaks as timing, amount, missing or duplicate, and route each class to an owner. Keep the provider’s identifier on every payment attempt and make the payment call itself idempotent so that a retry cannot charge twice. If you route billing through newer payment rails, such as the automated purchases described in agentic payments architecture, the same discipline applies: every attempt has a stable key and a reconciled outcome.
One more consideration is message format. If your invoices or remittance data flow to bank partners, the structured data standard matters; the ISO 20022 migration changes how rich payment references can be, which affects how easily incoming payments can be matched to invoices automatically.
Control totals as a product
A mature team publishes a daily billing health report: events accepted, duplicates dropped, late events by outcome, rejected events by error code, aggregate drift against raw, invoices drafted and finalised, and unreconciled payments. Stripe, for instance, emits error-report events for invalid meter events with codes such as a timestamp too far in the past or a missing customer, which is the right pattern: invalid data should produce a visible, countable signal and not disappear. Do the same in your own pipeline, and make dropped-event counts a first-class metric with an alert.
Customer-facing transparency
Usage-based invoices fail commercially when customers cannot see what they are paying for. Expose a usage explorer that matches the invoice: the same aggregates, the same windows, a drill-down to events or hourly buckets, and a near-real-time projection of the upcoming invoice. Stripe notes that aggregated usage in meter summaries and upcoming invoices might lag recent events because processing is asynchronous, so label any real-time view as an estimate and say when it was computed. Spend caps, threshold alerts and budget notifications turn a surprise invoice into a managed experience, and they are cheaper than support tickets.
Build Versus Buy: A Decision Matrix
The choice is not binary. Most teams end up with a hybrid: a vendor or open-source engine for invoicing and collection, and a thin in-house metering layer that owns the event contract. The table summarises the trade-offs; the vendor characterisations come from the project documentation I verified, and anything else is a general pattern rather than a claim about a specific product.
| Dimension | Payment-platform metering (for example Stripe Billing meters) | Open-source engine (for example Lago, OpenMeter) | Build in-house |
|---|---|---|---|
| Time to first invoice | Days; metering lives beside payments | Weeks; you deploy and operate it | Months |
| Pricing flexibility | Bounded by the platform’s models and limits | High; you can read and extend the code | Unlimited, and you pay for every case |
| Event throughput | Documented limits, such as 1,000 calls per second on the standard endpoint and 10,000 per second on the v2 stream | Scales with your infrastructure; OpenMeter is designed around Kafka and ClickHouse | Whatever you engineer |
| Idempotency and dedup | Identifier per event; semantics defined by the vendor | Varies; verify per project | Fully yours, fully your risk |
| Auditability and replay | Limited to what the vendor exposes | Depends on retained raw events | Strongest, if you design for it |
| Operational burden | Lowest | Medium to high | Highest |
| Licence and lock-in | Proprietary; migration requires data export | Lago is AGPLv3, OpenMeter is Apache 2.0; copyleft terms matter if you modify and offer it as a service | None |
| Best fit | Early-stage, simple tiers, Stripe already in use | Teams wanting control without writing the engine | Unusual pricing or strict data-residency needs |
Commercial specialists such as Metronome also exist in this space. I did not verify their current specifications for this article, so evaluate them against the same rows as above rather than relying on my summary.
A useful heuristic is to own the parts that encode your business and rent the parts that encode regulation. The event contract, the product’s definition of billable units and the price book are your business. Tax calculation, invoice numbering rules, payment-method handling and dunning are regulatory and operational plumbing where rented expertise pays off. Whatever you pick, insist on a raw-event export path. If you cannot get your own usage data out in a documented format, you do not own your revenue record.
AGPL deserves a separate note for legal review. It is a strong copyleft licence with obligations triggered by network use of modified versions. Many companies run it unmodified without issue, but the decision belongs to counsel, not to an engineering README.
Trade-offs, Gotchas, and What Goes Wrong
Double counting from retries. The classic. A client regenerates the event identifier per attempt, or two services emit the same business fact under different sources. Symptoms are bills that rise after a network incident. Defences: identifiers from the outbox, source-scoped keys, and a control-total check between emitter and ingestor.
Silent loss. The mirror image: an emitter fails open, a queue overflows, or an invalid event is dropped without a counter. Customers undercharged by a few per cent never complain, so you only find it in a revenue review. Defences: dead-letter queues with alerts, error-report metrics like the ones Stripe publishes for invalid meter events, and emitter heartbeats that state “I produced N events”.
Clock skew and timestamps. Producers stamping events with local clocks create future-dated events that land in the wrong window, and devices with wrong time rewrite history. Bound the accepted range, as Stripe does with its 35-day past and 5-minute future limits, and consider assigning ingestion time as a fallback when the producer clock is untrusted. Clock quality is a deep topic of its own, which is why reliable time sync matters for any distributed system.
Cardinality explosions. Letting customers or engineers attach arbitrary dimension values (a request identifier, a user identifier) multiplies the number of aggregate series. Stripe documents hard caps on unique dimension combinations for exactly this reason. Allow-list dimensions and review additions like schema changes.
Retroactive price changes. Re-rating history after a price change looks like a small job and turns into a mess of already-issued invoices. Enforce effective dating, and never re-rate a finalised period silently.
The cliff and the loophole. Volume tiers with cliffs invite gaming, such as splitting usage across accounts to stay under a threshold. Graduated tiers remove the cliff but can still be gamed by account fragmentation if tiers apply per account. Decide whether tiers apply per account, per organisation or per contract, and enforce it in the aggregation key.
Floating-point money. Using binary floats for currency introduces errors that appear in the fifth decimal and surface as one-cent mismatches against the provider. Use decimal types end to end, including in the database, and round once at a documented point.
Thundering herd at period end. Every customer’s invoice finalises at the same moment, rating spikes, and the payment provider throttles. Stagger anniversary billing, queue payment attempts with jitter and respect provider rate limits with exponential backoff, as Stripe recommends for 429 responses.
Replay side effects. A replay used to repair a rating bug must not re-send emails, re-charge cards or re-post ledger entries. Separate pure computation from side effects, and make side-effecting steps idempotent on their own keys.
Over-engineering. Not every product needs a streaming platform. If you bill a few hundred customers on a handful of metrics, a relational table with the schema above, a nightly aggregation and a human-reviewed draft invoice may be correct for years. Choose the simplest design that preserves the three invariants: append-only events, pure rating and immutable invoices.
Practical Recommendations
Start from the invariants, because they survive any technology choice. Events are immutable facts with producer-generated identifiers. Rating is a deterministic function of events and a versioned price book. Invoices are immutable documents with correction by new documents. If a proposed design violates one of these, it will eventually cost you a dispute.
Then sequence the work by risk. First, define the event contract and build the outbox plus idempotent ingestion, because errors here are unrecoverable once usage is lost. Second, add raw-event retention and a replayable aggregation, because that is your insurance. Third, build rating with golden-invoice tests drawn from real contracts. Fourth, add the late-data policy and the draft-to-final workflow. Last, add reconciliation and the customer usage explorer. Shadow-run the new system beside the old one for at least one full cycle and diff every invoice line before cutting over.
Checklist before going live:
- Every emitter derives its event identifier from the business action, never from the send attempt.
- Ingestion is insert-if-absent and returns success for duplicates.
- Timestamp bounds are enforced, and rejected events produce a countable signal.
- The dedup window exceeds the longest producer retry horizon, with the residual risk documented.
- A written late-event policy exists per customer class, with a grace period set from measured lag.
- Rating is pure, uses decimals and records the price book version on each line.
- Credit grants, drawdown and expiry are ledger entries, not mutable balances.
- Finalised invoices are database-enforced immutable; corrections use credit notes.
- Control totals compare emitters, raw events, rollups and provider settlements daily.
- A chaos replay with duplicates and resets reproduces identical aggregates.
- Raw events can be exported in a documented format.
This is educational systems analysis, not financial, tax, legal or accounting advice; confirm invoicing, tax and revenue-recognition obligations with qualified professionals.
Frequently Asked Questions
What is usage-based billing architecture?
Usage-based billing architecture is the set of systems that turn product consumption into invoices. It typically comprises event metering, idempotent ingestion into a durable log, aggregation by customer and window, a rating engine that applies a versioned price book, invoice generation, payment collection and reconciliation. Its defining requirement is that every billed quantity can be traced back to specific, deduplicated usage events and reproduced on demand.
How do you make usage event ingestion idempotent?
Have the producer generate a unique identifier at the moment the business action happens, usually through a transactional outbox, and reuse it on every retry. The ingestion service then performs an insert-if-absent on that identifier, scoped by tenant and source, and returns success for duplicates. This yields at-least-once delivery with an exactly-once effect on totals. The dedup retention window must be longer than the longest retry horizon.
How should late or out-of-order usage events be billed?
Define a watermark and a grace period from measured lag. If the invoice for the event’s window is still open, absorb the event and re-rate the draft. If it has been finalised, either carry the usage to the next invoice with its original event time, or issue a corrected or supplementary invoice for material amounts. Publish the policy, configure it per customer class, and never edit a finalised invoice.
What is the difference between graduated and volume pricing?
Graduated pricing charges each slice of usage at its own tier rate, so total cost always rises with usage. Volume pricing charges the entire quantity at the rate of the tier it lands in, which can create cliffs where slightly more usage costs less overall. With illustrative tiers of 0.10, 0.08 and 0.05 per unit, 12,500 units cost 945.00 under graduated pricing and 625.00 under volume pricing.
Should I build or buy a billing and metering system?
Own what encodes your business: the event contract, billable-unit definitions and the price book. Rent what encodes regulation and plumbing: tax, invoice numbering, payment methods and dunning. Payment-platform metering is fastest for simple tiers, open-source engines such as Lago or OpenMeter give control at an operational cost, and a full build suits unusual pricing. Whatever you choose, require a raw-event export so you retain your revenue record.
How do you reconcile usage-based invoices?
Reconcile in three layers. Compare emitter control totals with accepted events, compare rollups with raw events for the same window, and compare invoiced and collected amounts with the payment provider’s settlement data. Classify mismatches as timing, amount, missing or duplicate, and assign each class an owner. Store the inputs snapshot with each invoice so any dispute can be recomputed from the evidence.
Further Reading
- Idempotent API design: an engineering guide for key design and replay windows.
- Reconciliation engine architecture for payments for the cash-integrity layer.
- Agentic payments architecture for AI commerce for machine-initiated payment flows.
- ISO 20022 migration and payments architecture for structured payment references.
- Stripe documentation: record usage for billing, the primary reference for meter events, identifiers, timestamp limits and rate limits.
- CloudEvents specification v1.0.2, the source for the
idplussourceuniqueness rule.
By Riju — about
