Feature Flag Architecture: OpenFeature, Targeting Rules and Progressive Delivery at Scale
A feature flag looks like the simplest construct in software: an if statement whose condition lives somewhere else. That innocence is the trap. The moment the condition lives somewhere else, you have built a distributed configuration system that sits on the hot path of every request, decides what each user sees, and can take production down faster than any deploy. Most outages blamed on “a bad flag” are really failures of feature flag architecture: no safe default, no local cache, a rule engine that nobody tested at the edges, or a flag that outlived its purpose by three years.
This matters more in 2026 because release velocity has outrun deploy safety. Teams ship many times a day, and the only practical way to decouple deploying code from exposing behaviour is to put the exposure decision behind a runtime switch. Done well, flags give you kill switches, percentage rollouts, and experiments. Done badly, they give you an untestable combinatorial mess.
This article lays out a reference architecture, explains how evaluation, targeting and deterministic bucketing really work, and shows how OpenFeature standardises the application-facing API.
What this covers: the four flag types and their different lifetimes, server-side versus client-side evaluation, streamed local rules versus remote calls, hash bucketing math with a runnable Python sketch, the OpenFeature model (providers, hooks, evaluation context), flagd, outage behaviour, testing and flag debt, and a build-versus-buy matrix.
Context and Background
Feature toggles are an old idea. The vocabulary most teams use today was popularised by Pete Hodgson’s widely cited essay on feature toggles on martinfowler.com, which separates toggles by how long they live and how dynamic they must be. That framing is still the best starting point, because it explains why a single “flag service” cannot serve every use case equally well.
The commercial market grew around that need. Platforms such as LaunchDarkly, Unleash, Flagsmith and GrowthBook each offer a management UI, a rule engine, SDKs and some audit trail. Their internals differ, and I will not make claims about any vendor’s private architecture here. What they share, at the level of public documentation, is a split between a control plane where humans edit rules and a data plane where SDKs evaluate them.
The shift that changed the landscape was standardisation of the application-side API. Before OpenFeature, switching vendors meant rewriting every call site, because each vendor SDK exposed its own client, context type, and error model. OpenFeature defines a vendor-neutral specification for flag evaluation. The Cloud Native Computing Foundation announced it as an incubating project in December 2023, and you should check the CNCF project list for its current maturity level before citing it in a procurement document.
Progressive delivery is the other half of the story. Tools such as Argo Rollouts shift traffic between two versions of a service at the network layer. Flags shift exposure between two code paths inside one version. The two are complementary, and we compare them in detail later; if you run rollouts on edge fleets, our guide to Argo Rollouts progressive delivery for edge fleets covers the deployment-layer side.
One more framing before the architecture. A flag is an interface between three audiences with different needs: the engineer who writes the guarded code, the operator or product owner who flips it, and the analyst who reads the results. Good architecture serves all three without letting any of them break the others. Engineers need a typed, testable API and a guaranteed default. Operators need a UI that is fast, audited and hard to misuse. Analysts need exposure events that say exactly which variant each user saw and why. A system that optimises for only one of those audiences tends to fail in the other two.
The Reference Architecture: Control Plane, Distribution, and Local Evaluation
A production feature flag architecture has three layers: a control plane that stores versioned rules and an audit log, a distribution layer that streams or caches those rules close to the application, and an evaluation layer inside the SDK that resolves each flag locally against an evaluation context. Evaluating locally keeps latency in microseconds and keeps the application alive when the control plane is down.

Figure 1: Reference feature flag architecture. Rules flow one way from the control plane to local caches; evaluation never needs a round trip.
The diagram shows the shape that nearly every mature system converges on. Rules are authored in a UI or API, written to a versioned store, and recorded in an audit log. A distribution mechanism, either a streaming connection or a polled CDN object, pushes the rule set to server SDKs, edge workers and a relay for browsers. The application reads a flag value from a local in-memory structure. No request path should ever wait on the control plane.
Four flag types, four lifetimes
The most useful design decision is to refuse to treat all flags alike. Hodgson’s taxonomy gives four categories, and each has different requirements for dynamism and longevity.
Release flags hide unfinished or risky code behind a switch so it can be merged and deployed continuously. They should live for days or weeks and then be deleted. They are mostly static per environment, and the reason they exist is trunk-based development.
Ops flags control operational behaviour: disable an expensive recommendation call under load, shed a non-critical dependency, switch a read path to a replica. These are the kill switches. They may live for years, and they must be flippable in seconds by an on-call engineer who is not the original author.
Experiment flags assign users to variants for A/B and multivariate tests. They require stable, deterministic assignment, because a user who flips between variants mid-experiment corrupts the data. They live for the length of the test and are then deleted or converted.
Permission flags gate features by entitlement: a plan tier, a beta programme, a customer contract. They are long-lived, they overlap with authorisation, and they are the ones most likely to be mistaken for something that belongs in a proper entitlement service. My opinion: if a flag has been a permission flag for more than a year and encodes a pricing rule, promote it to a first-class entitlement model and remove it from the flag system.
Separate these categories in naming, in ownership metadata, and in expiry policy. A release flag should be created with an expiry date. An ops flag should be created with a runbook link. If your tooling cannot express that difference, you will end up with several hundred flags nobody dares to touch.
Server-side versus client-side evaluation
Where the rule engine runs determines your security and privacy posture.
In server-side evaluation, the SDK runs inside a trusted service. It can receive the full rule set, including targeting rules that mention customer identifiers, because the code is not exposed to end users. This is the cheapest and fastest model, and the one to prefer for anything sensitive.
In client-side evaluation, the code runs in a browser or mobile app that the user controls. You must not ship the raw rule set, because it would reveal targeting logic, segment membership and unreleased feature names to anyone who opens developer tools. The usual pattern is for the client to send its evaluation context to a trusted service, which evaluates all flags for that context and returns only the resulting values. The client caches those values and may refresh them on a stream.
The consequence is a real asymmetry. Server SDKs hold rules and evaluate per call. Client SDKs hold precomputed values and evaluate per context change. When the user’s context changes, say they log in, the client must re-fetch. That is why OpenFeature distinguishes a static-context paradigm for clients from a dynamic-context paradigm for servers, a distinction we return to below.
Local evaluation with streamed rules versus remote evaluation
There are two ways a server SDK can answer “what value does this flag have for this user”.
In remote evaluation, every call sends the flag key and context to a service and waits for the answer. It is simple, it keeps rules centralised, and it works with SDKs that have no rule engine at all. It also adds a network round trip to every evaluation, couples your availability to the flag service, and shifts your privacy surface because every context leaves your process.
In local evaluation, the SDK downloads the whole rule set, keeps it current through a stream or polling, and evaluates in-process. A call becomes a function over memory, which typically costs microseconds rather than milliseconds. The price is that the SDK must embed a rule engine whose behaviour matches the reference implementation exactly, and that rule sets must be small enough to ship to every instance.
I recommend local evaluation for server workloads in almost every case. The failure mode is the clincher: if the control plane disappears, a locally evaluating service keeps serving the last known rules, whereas a remotely evaluating service must fall back to defaults on every call. The OpenFeature Remote Evaluation Protocol (OFREP) exists for the cases where remote evaluation is genuinely needed, such as thin clients and languages without a mature in-process engine.
The long-description paragraph for Figure 1: left to right, authoring writes to a versioned flag store with an audit log, the store feeds a stream or CDN, and three consumers sit downstream: server SDKs with a local cache, edge workers with their own cache, and a relay proxy that serves browser clients precomputed values. Application code at the bottom only ever reads a resolved value.
Deeper Analysis: OpenFeature, Targeting and Deterministic Bucketing
What OpenFeature standardises
OpenFeature is a specification, not a flag service. Its specification is organised into the Flag Evaluation API, Providers, Evaluation Context, Hooks, Events and Tracking, with appendices covering the Remote Evaluation Protocol, observability and migrations. Requirements are written as normative statements using RFC 2119 keywords, so an implementation is compliant only if it meets all of them. Each section also carries a maturity label: experimental, hardening or stable. That label is worth reading before you depend on a feature, since only stable sections promise no breaking changes without a major version.
Four concepts do most of the work.
The client is what application code calls. It exposes typed methods such as boolean, string, number and object evaluation, each taking a flag key, a default value and optionally a context. Both a plain value and a detailed form are available; the detailed form returns the variant, the reason and any error code alongside the value.
The provider is the adapter that talks to a specific backend. It implements resolution for each type. Swapping vendors means swapping the provider registered at startup, not rewriting call sites. This is the portability promise, and it holds as long as you keep vendor-specific features out of your call sites.
The evaluation context is the bag of attributes about who or what is being evaluated: a targeting key, usually a user or session identifier, plus arbitrary attributes. Context can be set at the API level for the whole process, at the client level, and at the invocation level, and the layers merge with the more specific one winning.
Hooks are lifecycle callbacks around evaluation. This is where cross-cutting behaviour lives, so it is worth dwelling on.

Figure 2: OpenFeature evaluation sequence. Before hooks run API, client, then invocation; after, error and finally hooks run in the reverse direction.
The hook lifecycle in detail
The specification defines four hook stages. The before stage runs immediately ahead of resolution and, in the dynamic-context paradigm, may return additional evaluation context that is merged into the existing context for later hooks and the provider. The after stage runs following successful resolution and receives the evaluation details. The error stage runs when something fails in before, after or resolution itself. The finally stage runs unconditionally at the end.
Ordering matters and is stack-like. Before hooks run from API level to client to invocation to provider, in the order added within each level. After, error and finally hooks run in the opposite direction, with hooks within a level in reverse order of addition. This mirrors middleware onion models, and it lets a tracing hook open a span in before and close it in finally with the guarantee that outer wrappers close last.
Error semantics are deliberately forgiving. If a before hook raises, the remaining before hooks are skipped, error hooks run, and the caller receives the default value. Exceptions thrown inside error or finally hooks must be caught and must not propagate. Hook hints, a bag of string-keyed arbitrary properties passed through evaluation options, give per-call configuration to hooks and are immutable.
Practical uses of hooks include attaching enrichment such as the current tenant to the context, emitting an exposure event for experiment analysis, logging every evaluation with its reason, and recording a metric per flag. The last two are the foundation of flag observability, which I discuss under gotchas. The point is that these concerns are written once, in one hook, instead of at every call site.
Reasons and error codes are the debugging interface
Every detailed evaluation returns a reason. The specification lists STATIC, DEFAULT, TARGETING_MATCH, SPLIT, CACHED, DISABLED, UNKNOWN, STALE and ERROR, and notes that the field is not limited to these values. Reading them is a skill. TARGETING_MATCH means a rule matched explicitly. SPLIT means pseudorandom assignment decided the outcome. DEFAULT means the flag fell back to its configured default rather than a rule. STALE, a value that may be out of date, is the reason you want to alert on, because it means a provider is serving cached data it cannot confirm.
Errors use a closed set of codes: PROVIDER_NOT_READY, FLAG_NOT_FOUND, PARSE_ERROR, TYPE_MISMATCH, TARGETING_KEY_MISSING, INVALID_CONTEXT, PROVIDER_FATAL and GENERAL. Two of these deserve operational attention. TARGETING_KEY_MISSING means a percentage rollout was attempted without a stable identifier, and the result will be a default, silently diluting your rollout. PROVIDER_NOT_READY means code evaluated a flag before the provider finished initialising, a classic startup race that returns defaults for the first few seconds of a process lifetime.
Targeting rules: segments, attributes, and ordering
A targeting rule is a predicate over the evaluation context that selects a variant. Most engines support equality and set membership on attributes, comparisons, string matching and semantic-version comparison, then combine them with and, or and not. Rules are ordered, and the first match wins, with a fallthrough to a default rule at the end.
Two design points separate maintainable rule sets from tangled ones. First, prefer segments as named, reusable attribute predicates, such as “internal staff” or “enterprise tier”, over repeating the same conditions in dozens of flags. A segment is a single place to fix a mistake. Second, keep the rule count per flag small. A flag with twenty rules is a policy engine in disguise; each additional rule multiplies the number of cases your tests must cover, and nobody can reason about ordering past about five.
The rule language itself is a risk. A rich expression language lets engineers encode business logic in a UI no one code-reviews. The flagd project, for example, expresses targeting in JsonLogic, a JSON-based rule format, and extends it with custom operators. That is expressive and portable, and it also means rules are data that need their own tests and review process.
Deterministic hash bucketing for percentage rollouts
A percentage rollout needs a function that maps a user to a number between 0 and 100 such that it is uniform across users, stable across time and independent between flags. Randomness at evaluation time fails the stability requirement, because a user would see the feature on one request and not the next. Storing an assignment per user fails at scale. The standard solution is to hash.

Figure 3: Deterministic bucketing. Concatenate a stable key with the flag key, hash, reduce modulo a fixed number of buckets, compare with the rollout threshold.
The algorithm has five steps. Build a string from the targeting key, the flag key and optionally a per-experiment salt. Hash it with a fast, well-distributed function. Reduce the hash to an integer in a fixed range. Compare it with the threshold for the rollout percentage. Return the variant. Because the output depends only on the inputs, the same user gets the same answer on every server, every restart and every language that implements the same algorithm.
Including the flag key in the hash input is not optional. If you hashed only the user identifier, the same 5 percent of users would be selected for every rollout, forever. Those users would absorb every experimental risk and every experiment would be correlated. flagd’s fractional operator documents this explicitly: it hashes the bucketing value with murmur3 and recommends combining the flag key with a stable attribute such as an email address or session identifier, so that experiments on different flags are statistically independent. If you omit the bucketing expression, flagd concatenates the targeting key with the flag key.
Here is a self-contained Python sketch of the technique. It uses SHA-1 from the standard library for portability; production engines choose their own hash, and murmur3 is common, so do not expect assignments to match across engines.
import hashlib
BUCKETS = 10_000 # 0.01 percent resolution
def bucket(flag_key: str, targeting_key: str, salt: str = "") -> int:
"""Map a (flag, user) pair to a stable integer in [0, 9999]."""
raw = f"{salt}:{flag_key}:{targeting_key}".encode()
digest = hashlib.sha1(raw).digest()
return int.from_bytes(digest[:8], "big") % BUCKETS
def in_rollout(flag_key: str, user_id: str, percent: float) -> bool:
"""True when the user falls under the rollout threshold."""
return bucket(flag_key, user_id) < int(percent * 100)
def pick_variant(flag_key, user_id, weights):
"""weights: list of (variant, relative_weight). Sticky multivariate split."""
total = sum(w for _, w in weights)
point = bucket(flag_key, user_id) * total / BUCKETS
running = 0
for variant, weight in weights:
running += weight
if point < running:
return variant
return weights[-1][0]
I ran this against 200,000 synthetic user identifiers. At a 10 percent threshold, 10.03 percent of them landed in the rollout, close to the expected 10 percent. Raising the threshold from 5 to 10 percent kept every user who was already in, which is the monotonic property you want: ramping up never removes anyone. And two different flags at 10 percent each overlapped on 0.98 percent of users, which is what independence predicts (10 percent times 10 percent equals 1 percent). These are results from my own synthetic run, not a benchmark of any product.
The monotonic property deserves emphasis because it is easy to break. If a rollout is implemented as “bucket modulo 100 equals a randomly chosen residue”, increasing the percentage reshuffles who is in. Always implement rollout as “bucket below threshold” so that the set at 5 percent is a subset of the set at 10 percent.
Weights, sticky multivariate splits, and a caveat
flagd’s fractional operator generalises this to multiple variants using relative weights rather than percentages that must sum to 100, with a documented maximum weight sum of the 32-bit integer maximum, 2,147,483,647, which allows very fine granularity. Assignments are sticky: the same input yields the same variant unless the input or the configuration changes. That last clause is the caveat. Changing weights moves bucket boundaries, so users near a boundary will switch variants. For experiments, freeze the configuration for the duration of the test, or ramp only the treatment arm so that membership grows monotonically.
flagd and the in-process path
flagd is described by its maintainers as a feature flag evaluation engine and an OpenFeature-compliant backend. It has no UI, management console or persistence layer by design, and it is configured through a POSIX-style command line. It can read flag definitions from files, HTTP endpoints, Kubernetes custom resources and gRPC services, aggregate several sources, and apply changes as it observes them. It exposes flags over gRPC streams for in-process providers, over OFREP, and can run its evaluation engine inside the application process.
That makes flagd a good fit for two situations: teams that want a GitOps-style flag store with definitions reviewed as pull requests, and platform groups that want a reference implementation to test provider conformance against. It is a poor fit if you need product-manager self-service, experiment analysis or fine-grained audit, because it deliberately provides none of those. Treat it as the data plane, and bring your own control plane.
A minimal OpenFeature sketch in Python
The following shows the shape of the API using the OpenFeature Python SDK. Treat it as a sketch: package names, provider classes and signatures change between SDK versions, so check the current SDK documentation before copying it. The hook and context usage follows the specification described above.
from openfeature import api
from openfeature.evaluation_context import EvaluationContext
from openfeature.hook import Hook
class AuditHook(Hook):
def after(self, hook_context, details, hints):
print("flag", hook_context.flag_key,
"variant", details.variant, "reason", details.reason)
def error(self, hook_context, exception, hints):
print("flag error", hook_context.flag_key, repr(exception))
# api.set_provider(<provider for your backend, for example a flagd provider>)
api.add_hooks([AuditHook()])
client = api.get_client("checkout-service")
ctx = EvaluationContext(
targeting_key="user-4821",
attributes={"plan": "enterprise", "region": "eu"},
)
use_new_flow = client.get_boolean_value("new-checkout", False, ctx)
if use_new_flow:
... # new code path
else:
... # old code path
Note the second argument. The default value is False, which means that if the provider is unavailable, the targeting key is missing, or the flag does not exist, the old path runs. That default is a safety decision, and the next section treats it as one.
Outage behaviour: defaults are an architectural decision
Every flag evaluation has three outcomes: a rule result, a configured default inside the flag system, and the call-site default passed by the application. The call-site default is the last line of defence when the SDK cannot reach any state at all, for example on cold start before the first rule download.
Pick call-site defaults by asking what happens if this flag is unreachable at 3 a.m. For a release flag guarding new code, default to off, so an outage reverts to the proven path. For an ops kill switch that disables a risky dependency, the correct default is often the opposite of what the flag name suggests, because the safe state is the degraded one. For a permission flag, default to the entitlement a user is certain to have, never to the premium one, unless you want a flag-service outage to give away paid features.
Then layer the resilience. Persist the last good rule set to local disk so a restarted process can serve rules before it reconnects. Set a startup timeout, then proceed with defaults and log loudly rather than blocking boot indefinitely. Alert on the STALE reason and on PROVIDER_NOT_READY errors. Finally, rehearse the failure: block the flag service in a staging environment and confirm the application behaves as designed. Most teams discover their defaults are wrong in production.
Build versus buy
| Criterion | Build in-house | Open-source self-host | Commercial SaaS |
|---|---|---|---|
| Time to first flag | Weeks | Days | Hours |
| Control-plane UI and RBAC | You build it | Varies by project | Included |
| Audit trail and approvals | You build it | Varies | Typically included, verify tier |
| Experiment statistics | Rarely worth building | Varies | Often included, verify |
| Data residency and control | Full | Full | Contractual, verify |
| Ongoing operating cost | Engineering time | Infra plus engineering time | Subscription |
| Lock-in risk | Low | Low to medium | Mitigated by OpenFeature providers |
| Best fit | Tiny flag sets, special constraints | Platform teams with Kubernetes skills | Product-led teams needing self-service |
The honest guidance: build only the thing the market cannot sell you. A flag evaluation engine with consistent semantics across languages, a streaming protocol and an audit UI is multi-year work, and the interesting engineering in your company is elsewhere. A thin wrapper over an open source engine is reasonable. Writing your own bucketing hash is not. Whatever you choose, route all call sites through OpenFeature so that the decision is reversible, and verify current pricing, tier limits and compliance features directly with each vendor, since I have not verified them for this article.
Trade-offs, Gotchas, and What Goes Wrong
Flags versus canary deployments
Flags and canary releases are often presented as alternatives. They are different tools with different blast radii. A canary or progressive rollout at the deployment layer exposes a new build to a slice of traffic and watches infrastructure and service metrics; it protects you from bad binaries, bad configuration and resource regressions. A flag exposes a code path inside a build that is already running everywhere; it protects you from bad behaviour while leaving the binary unchanged. If you want the detail on the deployment-layer tools, our Argo Rollouts versus Flagger decision record compares them.
The strongest pattern combines both. Deploy the new build by canary with the new code path flagged off, confirm the build is healthy, then ramp the flag independently. This separates two risks that are otherwise tangled: “does this binary run” and “do users like this behaviour”. When something breaks, you know which lever to pull, and rolling back a flag takes seconds where rolling back a deployment takes minutes.
The testing matrix explodes
Every boolean flag doubles the number of reachable program states. Ten independent flags imply 1,024 combinations, and no test suite covers them. In practice you do not need to, because most flags are independent, but “most” is doing a lot of work in that sentence. The dangerous ones are interactions: two flags that both change how an order total is computed, or a release flag nested inside a permission flag.
Contain this with discipline rather than heroics. Test each flag’s two states in isolation, and test the production combination (all flags at their current values) plus the combination you intend to ship next. Forbid nested flag checks where one flag’s meaning depends on another. Make flags readable from tests through a test provider, which OpenFeature’s provider model supports well, so a test can pin a flag without touching a network. Finally, record the full set of evaluated flags and variants in your logs and traces, so that when a bug reproduces only for one user you can recreate their exact flag combination.
Flag debt
Flag debt is the accumulation of flags that have served their purpose and remain. It is insidious because each one is harmless alone. Collectively they make the code harder to read, since every conditional is a fork, they slow builds, they expand the attack surface, and they create a nightmare of unknowns when you try to remove them years later. Worse, a stale flag is a loaded gun: someone eventually flips a seven-year-old switch and discovers the dead branch it guards no longer works.

Figure 4: Flag lifecycle. Every release flag is created with an owner and expiry and ends in deletion; the kill switch loops back to dark launch.
The fix is lifecycle management enforced by tooling, not by good intentions. Require an owner and an expiry date at creation. Report flags past expiry in a weekly list, and raise a ticket automatically. When a flag has been at 100 percent for a soak period, such as a couple of weeks, delete the flag and the dead branch together in one change. Treat removal as part of the definition of done for the feature itself.
Some teams add a lint rule that fails the build when a release flag is past its expiry date. That sounds heavy-handed until the first time it prevents a forgotten flag from shipping into a new service. The cost of keeping a flag compounds, and the cheapest day to delete it is the day you are confident in the feature.
Security, privacy and audit
A flag change is a production change. It can alter pricing, disable fraud checks, or expose unreleased features, yet in many organisations it bypasses the review and approval a code change would receive. Require role-based access, restrict who can change production flags, and require approvals for high-risk flags. Keep an immutable audit log with who changed what, when, from which value to which value. That log is also your best incident-response tool: most flag-induced incidents are answered by the question “what changed in the last ten minutes”.
Privacy deserves a short note. The evaluation context often contains personal data. With local evaluation, that data stays in your process. With remote evaluation, it travels to a third party. Send only the attributes that rules actually need, prefer opaque identifiers over email addresses, and check your data-processing agreements before sending context to a vendor.
Performance and consistency pitfalls
Local evaluation is cheap, but not free. A rule set with thousands of flags and complex regular expressions costs memory per instance and can add measurable latency on a hot path, so measure rather than assume. Consistency is a second concern: two instances may hold slightly different rule versions during a propagation window, so a single user request that passes through two services might see two different values for the same flag. For flags whose value must be consistent within a request, evaluate once at the edge and propagate the resolved value, for example in a request header or in baggage on a distributed trace, rather than re-evaluating at every hop.
Finally, avoid evaluating flags in tight loops. Resolve once per request, store the value in the request scope, and read from there. Evaluation hooks that log every call turn a cheap lookup into a firehose of log lines; sample them or log only on change.
Practical Recommendations
Start with the lifecycle, not the tool. Decide the four flag types, give each a naming convention and a maximum lifetime, and encode expiry as required metadata. This single step prevents most of the long-term pain.
Standardise on OpenFeature at the call site, even if you use a single vendor today. The cost is one thin provider layer; the benefit is that you can change backend, add a local flagd for testing, or run a second backend during migration without touching application code. Check the maturity label on any specification section you rely on, and pin SDK versions.
Evaluate locally on servers, and keep rules out of browsers. For ops flags, choose call-site defaults as an explicit safety decision and review them like you would review a circuit breaker setting. Persist the last good rule set and alert on stale evaluations.
Make rollouts monotonic and independent by hashing the flag key together with a stable identifier, comparing against a threshold, and never reshuffling users when you ramp. Log which variant each user saw and why, as that is the data both your experiments and your incident reviews depend on.
Pair flags with deployment-layer progressive delivery rather than choosing between them, and write down which lever handles which kind of risk.
A short checklist you can adopt this week:
- Every flag has an owner, a type and an expiry date at creation.
- Call-site defaults are reviewed and documented for each ops and permission flag.
- Server SDKs evaluate locally and persist the last good rule set.
- Percentage rollouts use flag key plus stable identifier, with a threshold comparison.
- Production flag changes require role-based access and leave an audit record.
- A weekly report lists flags past expiry; removal is part of the definition of done.
- The failure of the flag service is rehearsed in staging at least once.
Frequently Asked Questions
What is feature flag architecture?
Feature flag architecture is the design of the system that decides, at runtime, which code path a given user or request takes. It has a control plane where rules are authored and audited, a distribution layer that delivers rules to applications, and an evaluation layer inside the SDK that resolves each flag against an evaluation context. A good design evaluates locally, has safe defaults, and treats flag lifecycle as a first-class concern.
What is OpenFeature and is it a feature flag service?
OpenFeature is a vendor-neutral specification and set of SDKs for flag evaluation, not a flag service. It defines a client API, providers that adapt to specific backends, an evaluation context, hooks, events and tracking. Your code calls the OpenFeature client, and a provider talks to your chosen backend. It was announced as a CNCF incubating project in December 2023, and its current maturity should be confirmed on the CNCF project list.
How does percentage rollout bucketing work?
The system hashes a stable identifier together with the flag key, reduces the result to an integer in a fixed range such as 0 to 9,999, and compares it with a threshold derived from the rollout percentage. Because the hash is deterministic, a user receives the same result every time. Including the flag key makes different rollouts statistically independent, and a threshold comparison guarantees that ramping up never removes users who were already enabled.
Should flags be evaluated on the server or the client?
Evaluate on the server whenever possible, because the rule set can include sensitive targeting logic that must not be exposed to users. For browsers and mobile apps, have a trusted service evaluate for the user context and return only the resulting values, which the client caches and refreshes. OpenFeature models this as a static-context paradigm for clients and a dynamic-context paradigm for servers, with context changes triggering re-evaluation on the client.
What is flag debt and how do you avoid it?
Flag debt is the build-up of flags that have finished their job but remain in code and configuration. They add branches, slow comprehension, and become hazardous when someone flips them years later. Avoid it by requiring an owner and expiry at creation, reporting flags past expiry, deleting the flag and the dead code together once a feature is stable at 100 percent, and making removal part of the definition of done.
Are feature flags a replacement for canary deployments?
No. Canary and progressive deployment tools shift traffic between builds and protect you from bad binaries or infrastructure regressions. Flags change behaviour within one running build and protect you from bad features. They are complementary: deploy by canary with the feature flagged off, confirm the build is healthy, then ramp the flag separately. Using both keeps two different risks independently reversible, one in minutes and one in seconds.
Further Reading
- Argo Rollouts progressive delivery for edge fleets: the deployment-layer counterpart to flags.
- Argo Rollouts versus Flagger: a progressive delivery decision record: how to choose a rollout controller.
- 3GPP Release 20 and 5G-Advanced RedCap for industrial IoT: a connectivity context where staged rollouts to constrained devices matter.
- 5G RedCap for industrial IoT: reduced capability in 3GPP Release 17 and 18: device-fleet background for remote configuration.
- OpenFeature specification: the normative reference for the API, providers, context and hooks.
- flagd documentation: the evaluation engine and its fractional operator.
- Feature Toggles by Pete Hodgson: the classic taxonomy of toggle types.
By Riju — about
