Kubernetes ValidatingAdmissionPolicy and CEL: Admission Control Without Webhooks

Kubernetes ValidatingAdmissionPolicy and CEL: Admission Control Without Webhooks

Kubernetes ValidatingAdmissionPolicy and CEL: Admission Control Without Webhooks

Every cluster that enforces policy eventually learns the same lesson: the thing that guards your API server can also take it down. A validating admission webhook is a web service sitting in the request path of every matching write. When its pods are unscheduled, its certificate expires, or its latency creeps up, deployments stall, node drains hang, and in the worst case the webhook cannot restart because it is blocked by its own rule.

ValidatingAdmissionPolicy removes that failure class for the large family of rules that are pure functions of the request. The rule is written in the Common Expression Language (CEL), stored as an API object, and evaluated inside the kube-apiserver process. It went stable in Kubernetes v1.30, and its mutating sibling, MutatingAdmissionPolicy, is documented as stable since v1.36.

This guide walks the full model: the three resources, the CEL environment, parameters, validation actions, failure behaviour, testing, a safe rollout path, and an honest decision framework against webhooks, Kyverno, and Gatekeeper. You get runnable policies for replica limits, image tags, and required labels.

What this covers: the request path, the policy, binding, and parameter model, CEL variables and cost limits, three production-grade policies, local testing, MutatingAdmissionPolicy status, trade-offs, and a decision matrix.

Context and Background

Admission control is the last gate before an object is persisted. After authentication and authorization, a write request passes through mutating admission, then schema validation, then validating admission, and only then reaches etcd. The built-in admission plugins handle things like namespace lifecycle and resource quotas. For custom rules, Kubernetes historically offered exactly one extension point: dynamic admission webhooks, registered with ValidatingWebhookConfiguration and MutatingWebhookConfiguration.

Webhooks are flexible because they are arbitrary code. They can call a database, consult an external inventory, or run a policy engine with its own language. That flexibility has a price that the project itself documents. The Kubernetes admission webhook good practices guide tells operators to set small timeouts, load balance the webhook, run it highly available, exclude the namespace it lives in, and avoid dependency loops between webhooks and cluster add-ons. In its dependency-loop section it recommends using ValidatingAdmissionPolicies to avoid introducing dependencies at all.

The webhook configuration reference gives concrete numbers. A webhook timeout must be between 1 and 30 seconds and defaults to 10 seconds, and failurePolicy defaults to Fail, so an unreachable webhook rejects matching requests. The same documentation notes that mutating webhooks are called in sequence and so add latency one after another, while validating webhooks are called in parallel.

Policy engines such as Kyverno and OPA Gatekeeper wrap this webhook mechanism in a friendlier policy language, with reporting, exceptions, and libraries of ready-made rules. They are the right tool for many teams, and our comparison of Kyverno and OPA Gatekeeper covers that terrain. But they inherit the webhook availability model. Both projects have also been adding ways to lean on the native CEL engine, so the boundary between “engine” and “built-in” is moving.

ValidatingAdmissionPolicy changes the economics. The official Validating Admission Policy documentation describes it as “a declarative, in-process alternative to validating admission webhooks”. There is no server to deploy, no TLS to rotate, no service to keep alive, and no network hop. The cost is a deliberately limited language: CEL is non-Turing complete, side-effect free, and bounded by a deterministic cost budget. That limitation is the feature.

One more thing to hold in mind before the mechanics. Version facts matter here, and they have moved. The v1.30 GA claim for ValidatingAdmissionPolicy is confirmed by the current docs. The mutating counterpart reached stable in v1.36 according to its documentation page, and the docs site currently lists v1.37 as the newest release. If you run an older control plane, check your version before assuming any feature in the second half of this article.

The ValidatingAdmissionPolicy Model: Three Resources, One In-Process Evaluator

ValidatingAdmissionPolicy is a declarative rule evaluated inside the kube-apiserver by a CEL interpreter. A working policy needs at least two objects, the policy and a binding, and optionally a third, a parameter resource. The policy holds the logic, the binding scopes it and picks the enforcement action, and the parameter object supplies configuration values.

ValidatingAdmissionPolicy request path showing in-process CEL evaluation next to out-of-process admission webhook calls

Figure 1: Where admission policy runs in the API request path. Policies and mutating policies execute inside the apiserver, while webhooks add a network call per matching request.

Figure 1 places the policy engines in the request pipeline. The key structural point is the dashed edges: webhook servers sit outside the control plane process, and every matching request pays a network round trip and inherits that server’s availability. Admission policies sit on the solid path, so their latency is CPU time inside the apiserver and their availability is the apiserver’s own.

The policy: abstract logic

The ValidatingAdmissionPolicy object, in admissionregistration.k8s.io/v1, declares what it inspects and what must be true. The docs describe it as the abstract logic of a policy, for example “this policy makes sure a particular label is set to a particular value”. The important fields are these:

  • spec.matchConstraints lists resourceRules (API groups, versions, operations, resources) plus optional selectors. It decides which requests the policy even looks at.
  • spec.validations is a list of CEL expressions. Each must evaluate to true for the request to pass. Each can carry a static message, a dynamic messageExpression, and a machine-readable reason.
  • spec.matchConditions adds CEL-based pre-filters for cases that rules and selectors cannot express.
  • spec.variables names reusable sub-expressions.
  • spec.auditAnnotations emits key-value annotations into the audit log.
  • spec.paramKind declares the kind of the parameter resource, if any.
  • spec.failurePolicy decides what happens when the policy itself misbehaves.

The binding: scope and enforcement

A policy alone does nothing. The documentation is explicit that at least a policy and a corresponding ValidatingAdmissionPolicyBinding must exist for it to have an effect. The binding links a policy by policyName, narrows scope with matchResources (for example a namespaceSelector), optionally attaches a parameter through paramRef, and declares validationActions.

This split is the design’s best idea. One abstract policy can be bound many times: strict Deny in production namespaces, Warn and Audit in development, each with different parameter values. The author of the logic and the operator who decides where it applies can be different teams with different RBAC.

The parameter: configuration without redeploying logic

A parameter resource separates configuration from logic. It can be a ConfigMap or any custom resource. The policy declares paramKind (a group, version, and kind), and the binding points at an instance with paramRef, either by name and optional namespace, or by a label selector. Exactly one of name or selector must be set.

The semantics worth memorising come straight from the docs. For each admission request the API server evaluates the CEL expressions of each matching combination of policy, binding, and parameter, and the request must pass all of them. If several bindings match, the policy is evaluated once per binding. If a selector matches several parameter objects, the rules run for each and the results are ANDed. Overlap is allowed, and it multiplies evaluations.

Policy binding and parameter model: one ValidatingAdmissionPolicy bound twice with different namespace selectors, actions, and parameters

Figure 2: One policy, two bindings. Production namespaces get a strict parameter and Deny, development namespaces get a looser parameter and Warn plus Audit.

Figure 2 shows the fan-out. Each binding contributes its own scope, its own parameter, and its own action set, and all of them feed the same CEL evaluation. Because evaluations multiply with overlapping bindings, keep binding scopes disjoint unless you intend layered rules.

paramRef.parameterNotFoundAction is required, with values Allow or Deny. Allow treats missing parameters as a pass. Deny applies the policy’s failurePolicy, so with Fail the request is rejected. When a policy needs parameters, the docs recommend a guard as the first validation: params != null with a message such as “params missing but required to bind to this policy”. Without it, a typo in paramRef can silently disable a policy or break every request, depending on how you chose.

The CEL Environment: Variables, Optional Fields, and Cost

CEL is a small expression language from Google: C-like syntax, strong typing, no loops, no recursion, no side effects. Kubernetes uses it in CustomResourceDefinition validation rules and in admission policy. Because it is non-Turing complete, every expression terminates, and Kubernetes layers a deterministic cost model on top.

The variables you actually have

Validation expressions see these variables, per the current docs:

  • object: the incoming object; null for DELETE.
  • oldObject: the existing object; null for CREATE.
  • request: attributes of the admission request, including operation, userInfo, and resource.
  • params: the bound parameter resource; null when no paramKind is set.
  • namespaceObject: the namespace the object lives in; null for cluster-scoped objects.
  • authorizer: a CEL authorizer for checking the requesting principal’s permissions. authorizer.requestResource is a shortcut preconfigured with the request’s resource.
  • variables: the lazily evaluated map of your named sub-expressions.

object and oldObject are strongly typed against the resource schema, so object.spec.template.spec.containers is checked, not guessed. For any object, including schemaless custom resources, CEL guarantees access to only apiVersion, kind, metadata.name, and metadata.generateName.

Two scoping rules trip people up. Match conditions can use the same variables except namespaceObject and variables: the namespace object is not populated there and always evaluates to null, and variables are not yet defined because match conditions run first. To filter by namespace labels in a match, use namespaceSelector in matchConstraints or the binding’s matchResources. Message expressions, by contrast, can use object, oldObject, request, params, and namespaceObject.

Optional fields without crashing

CEL errors on absent fields, and an error is not the same as false. The docs give two safe patterns. Use has(object.field) for field presence, and the in operator for map keys: has(object.metadata.labels) && 'example.com/environment' in object.metadata.labels. For deep traversal, Kubernetes 1.29 and later supports CEL optional syntax: object.?metadata.labels['example.com/block'].orValue('') returns a default instead of failing on absent intermediates.

Boolean short-circuiting makes optional parameters tidy. The docs’ example !has(params.optionalNumber) || (params.optionalNumber >= 5 && params.optionalNumber <= 10) only checks the range when the parameter exists.

Variable composition

When a sub-expression is reused or expensive, move it into spec.variables. Variables are evaluated lazily on first reference, and both result and any error are memoised, so they count once toward cost. Order matters: a variable may reference only variables defined before it, which prevents cycles. The docs’ flagship example derives an environment variable from the namespace label and checks every container image against it, in a few lines.

Cost: the safety net

The Kubernetes CEL reference explains that cost units are deterministic and independent of hardware or load. Simple operations such as comparisons cost 1, a list literal has a fixed base cost of 40, and regular expression functions are estimated at length(regexString) * length(inputString), reflecting worst-case RE2 behaviour. Every expression runs under a runtime cost budget; exceed it and evaluation halts with an error.

That error is then handled by failurePolicy. I am deliberately not quoting numeric budget limits here: the page describes the mechanism and the specific limits are best read from the version of the docs matching your cluster. What the mechanism means in practice is simple. Regex over long strings inside all() over big lists is where cost accumulates, so measure before you ship a clever one-liner.

Writing and Operating Real Policies: A Walk-through

Theory is cheap, so here are three policies you can paste into a test cluster. Each one is a rule that organisations routinely implement as a webhook, and each fits the pure-function profile. Replica limits need the request and a number. Image tags need the request. Required labels need the request. None of them needs to call anything.

Policy 1: a parameterised replica limit

This policy uses a ConfigMap as the parameter, so no CustomResourceDefinition is required. The binding for production reads one ConfigMap and the binding for everything else reads another.

apiVersion: v1
kind: ConfigMap
metadata:
  name: replica-limit-prod
  namespace: policy-params
data:
  maxReplicas: "10"
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: replica-limit-other
  namespace: policy-params
data:
  maxReplicas: "50"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: replica-limit.policy.example.com
spec:
  failurePolicy: Fail
  paramKind:
    apiVersion: v1
    kind: ConfigMap
  matchConstraints:
    resourceRules:
    - apiGroups: ["apps"]
      apiVersions: ["v1"]
      operations: ["CREATE", "UPDATE"]
      resources: ["deployments"]
  validations:
  - expression: "params != null && has(params.data) && 'maxReplicas' in params.data"
    message: "replica-limit parameter ConfigMap is missing or has no maxReplicas key"
  - expression: "object.spec.replicas <= int(params.data.maxReplicas)"
    messageExpression: "'spec.replicas is ' + string(object.spec.replicas) + ' but this namespace allows at most ' + params.data.maxReplicas"
    reason: Invalid
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: replica-limit-prod.policy.example.com
spec:
  policyName: replica-limit.policy.example.com
  validationActions: [Deny]
  paramRef:
    name: replica-limit-prod
    namespace: policy-params
    parameterNotFoundAction: Deny
  matchResources:
    namespaceSelector:
      matchLabels:
        environment: prod
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: replica-limit-other.policy.example.com
spec:
  policyName: replica-limit.policy.example.com
  validationActions: [Warn, Audit]
  paramRef:
    name: replica-limit-other
    namespace: policy-params
    parameterNotFoundAction: Deny
  matchResources:
    namespaceSelector:
      matchExpressions:
      - key: environment
        operator: NotIn
        values: ["prod"]

Three details deserve attention. The first validation is the parameter guard recommended in the docs, adapted to a ConfigMap. The second uses messageExpression, which per the docs takes precedence over a static message but falls back to it if the expression fails, or if the result is a multi-line string. And reason: Invalid sets the HTTP response mapping; the documented reasons are Unauthorized, Forbidden, Invalid, and RequestEntityTooLarge, defaulting to Invalid when unset.

One caveat that bites: a namespace with no environment label matches the second binding through NotIn, which is usually what you want, but verify that behaviour against your own namespace conventions before relying on it. Also note that int() on a malformed string produces an evaluation error, which failurePolicy then governs.

Policy 2: no mutable image tags

The rule: every container, including init containers, must pin either a digest or an explicit non-latest tag. A bare nginx implies latest and is rejected. The regular expressions below were exercised locally (see the testing section) against registry-with-port images, which are the classic trap, because host:5000/img contains a colon but no tag.

apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: pinned-image-tags.policy.example.com
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
    - apiGroups: ["apps"]
      apiVersions: ["v1"]
      operations: ["CREATE", "UPDATE"]
      resources: ["deployments", "statefulsets", "daemonsets"]
  variables:
  - name: allContainers
    expression: >-
      object.spec.template.spec.containers +
      object.spec.template.spec.?initContainers.orValue([])
  validations:
  - expression: >-
      variables.allContainers.all(c,
        c.image.matches('@sha256:[a-f0-9]{64}$') ||
        (c.image.matches('^[^@]*[^/@]:[^/:@]+$') && !c.image.endsWith(':latest')))
    messageExpression: >-
      'images must use a digest or an explicit tag other than latest: ' +
      variables.allContainers.filter(c,
        !(c.image.matches('@sha256:[a-f0-9]{64}$') ||
          (c.image.matches('^[^@]*[^/@]:[^/:@]+$') && !c.image.endsWith(':latest')))
      ).map(c, c.image).join(', ')
    reason: Invalid
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: pinned-image-tags.policy.example.com
spec:
  policyName: pinned-image-tags.policy.example.com
  validationActions: [Warn, Audit]
  matchResources:
    namespaceSelector:
      matchExpressions:
      - key: kubernetes.io/metadata.name
        operator: NotIn
        values: ["kube-system"]

This policy has no paramKind, so params is null and the binding needs no paramRef. The variables entry builds the combined container list once, using the optional-field syntax so a Deployment without init containers does not raise an error. The binding starts in Warn and Audit mode, a point we return to when discussing rollout. It also excludes kube-system by name through the automatic kubernetes.io/metadata.name namespace label.

Notice the cost angle: matches() inside all() runs once per container, and the message expression repeats the predicate. Container lists are short, so this is fine. If you were validating thousands of list entries you would reach for a variable that computes the offending subset once.

Policy 3: required ownership labels, with transition rules

The third policy requires labels on creation and forbids removing them later, which demonstrates oldObject and request.operation.

apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: required-labels.policy.example.com
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
    - apiGroups: ["apps"]
      apiVersions: ["v1"]
      operations: ["CREATE", "UPDATE"]
      resources: ["deployments"]
  matchConditions:
  - name: skip-system-users
    expression: "!request.userInfo.username.startsWith('system:')"
  variables:
  - name: required
    expression: "['app.kubernetes.io/name', 'owner']"
  validations:
  - expression: >-
      has(object.metadata.labels) &&
      variables.required.all(k, k in object.metadata.labels)
    message: "deployments must carry the labels app.kubernetes.io/name and owner"
    reason: Invalid
  - expression: >-
      request.operation == 'CREATE' ||
      !has(oldObject.metadata.labels) ||
      !('owner' in oldObject.metadata.labels) ||
      object.metadata.labels['owner'] == oldObject.metadata.labels['owner']
    message: "the owner label is immutable once set"
    reason: Forbidden
  auditAnnotations:
  - key: "owner-label"
    valueExpression: "has(object.metadata.labels) && 'owner' in object.metadata.labels ? object.metadata.labels['owner'] : 'unset'"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: required-labels.policy.example.com
spec:
  policyName: required-labels.policy.example.com
  validationActions: [Deny, Audit]

Two things are worth explaining. The matchConditions entry skips any request whose username begins with system:, which exempts controllers such as the deployment controller operating on existing objects. Match condition failures follow a specific rule: if any condition evaluates to false the policy is skipped, and if a condition errors the outcome depends on failurePolicy (reject for Fail, skip for Ignore). Exempting by username prefix is a convenience, not a security boundary, so decide deliberately who should be exempt.

The second validation uses reason: Forbidden, so a client sees a 403 for the immutability violation and a 422 for the missing labels. The auditAnnotations entry writes a key prefixed with the policy name, here required-labels.policy.example.com/owner-label, into the audit event. The docs note that if the expression evaluates to null the annotation is omitted, and that on key collisions with other admission controllers the first one wins.

How one request flows through the engine

Sequence diagram of a request through ValidatingAdmissionPolicy with Deny, Warn and Audit, and pass outcomes

Figure 3: Evaluation outcomes for a single matched request. Deny rejects with a message and reason, Warn and Audit admit with a header and an audit annotation, and a clean pass is silent.

The sequence in Figure 3 makes the action semantics concrete. Per the docs, Deny rejects the request, Warn returns the failure to the client as an HTTP warning, and Audit records the failure in the audit event for the request. They combine, with one prohibition: Deny and Warn may not be used together, because the same failure would be reported twice, once in the response body and once in the warning header.

A validation that evaluates to false is always enforced according to the binding’s actions. A failure that comes from failurePolicy, meaning a misconfiguration or CEL runtime error, is enforced according to those actions only when failurePolicy is Fail; with Ignore it is dropped silently. That asymmetry matters for observability, which we revisit below.

Type checking: the free linter

When you create or update a policy, the apiserver parses every expression and rejects syntax errors outright. Afterwards it type-checks referenced variables against the types named in matchConstraints and writes results to status.typeChecking. A present-but-empty status.typeChecking means no errors were found. The docs’ example is a Deployment expression object.replicas > 1, which produces a warning that the field replicas is undefined, pointing at the character position.

The limitations are documented and important. Type checking never changes behaviour, so a policy with warnings still evaluates. Wildcards are not checked, so a policy matching "*" in group, version, or resource gets no type check on what the wildcard covers. At most 10 matched types are checked, and the eleventh and beyond are ignored. CRDs, including paramKind references, are not type checked. Always read status.typeChecking after applying a policy:

kubectl apply -f pinned-image-tags.yaml
kubectl get validatingadmissionpolicy pinned-image-tags.policy.example.com \
  -o jsonpath='{.status.typeChecking}{"\n"}'

Testing CEL locally before the cluster sees it

Server-side dry runs are the highest-fidelity test, but you can unit test expression logic faster with a CEL implementation. The community cel-python package (celpy) is a pure-Python CEL interpreter. It does not include Kubernetes’ extension libraries (the quantity, URL, or authorizer functions), so use it for core-language logic and regex behaviour, not for anything depending on the Kubernetes CEL library. The script below mirrors the tag rule above and was run in a sandbox before publishing:

# pip install cel-python
import celpy

env = celpy.Environment()

def evaluate(expr, ctx):
    return env.program(env.compile(expr)).evaluate(celpy.json_to_cel(ctx))

DIGEST = r"c.image.matches('@sha256:[a-f0-9]{64}$')"
TAG = r"c.image.matches('^[^@]*[^/@]:[^/:@]+$') && !c.image.endsWith(':latest')"
RULE = f"({DIGEST}) || ({TAG})"

CASES = {
    "nginx": False,                      # implicit latest
    "nginx:latest": False,
    "nginx:1.27": True,
    "host:5000/img": False,              # port, not a tag
    "host:5000/img:2.0": True,
    "repo/app@sha256:" + "a" * 64: True, # digest pin
}

for image, expected in CASES.items():
    got = bool(evaluate(RULE, {"c": {"image": image}}))
    assert got == expected, (image, got, expected)
    print(f"{image[:40]:<42} {got}")

# Required labels rule
LABELS = "['owner','app.kubernetes.io/name'].all(k, k in object.metadata.labels)"
assert not evaluate(LABELS, {"object": {"metadata": {"labels": {"owner": "x"}}}})
assert evaluate(LABELS, {"object": {"metadata": {"labels": {
    "owner": "x", "app.kubernetes.io/name": "api"}}}})
print("all assertions passed")

For cluster-level tests use the real thing. Create the policy with the binding set to [Warn] in a scratch namespace and run kubectl apply --dry-run=server -f deployment.yaml. A dry run goes through admission, so warnings and denials appear exactly as in a real write without persisting anything. The CEL Playground is also pointed to by the official docs for quick experiments. For a regression suite, keep a directory of “should pass” and “should fail” manifests and run both through server-side dry run in CI against a disposable cluster.

Safe rollout: Audit, then Warn, then Deny

Because actions live on the binding and not the policy, you can promote enforcement without touching the logic. The pattern that works:

  1. Create the policy and a binding with validationActions: [Audit]. Nobody sees anything. Mine the audit log for the policy’s annotation or failure records to learn how many existing workloads would break.
  2. Switch to [Warn]. Engineers see warnings from kubectl on their next apply. Fix the noisy teams’ manifests.
  3. Switch to [Deny] namespace by namespace, using several bindings with disjoint selectors, rather than flipping the whole cluster.

Remember that policies evaluate on UPDATE, so an existing non-compliant Deployment becomes un-editable the moment Deny lands, including for routine scale operations unless your rule is written to tolerate them. A transition rule that only checks fields which changed (compare object with oldObject) is a gentler way to introduce a constraint on live objects.

MutatingAdmissionPolicy: The Other Half

Validation says no. Mutation says “let me fix that for you”, such as injecting a default label, adding a sidecar, or setting a security context. Historically that meant a mutating webhook, with all the availability concerns above and an extra one: mutating webhooks are called in sequence and there is no stable order in which mutations are applied. The webhook good-practices page tells operators to add a validating admission step afterwards to confirm the mutation survived.

The MutatingAdmissionPolicy documentation states that the feature is stable since Kubernetes v1.36 and enabled by default, and that it was first available in v1.30. The resources are MutatingAdmissionPolicy and MutatingAdmissionPolicyBinding, shown in the docs under admissionregistration.k8s.io/v1. If your cluster is older than 1.36, treat it as a pre-stable feature and check which API version and feature gate your release uses, because I have not verified the intermediate alpha and beta history and you should not rely on it from this article.

Two ways to express a mutation

A mutating policy has a mutations list, which cannot be empty. Each entry has a patchType:

  • ApplyConfiguration: a CEL expression that builds an object using the Object{} initialiser. The result is merged into the live object using the server-side apply merge strategy. Apply configurations may not modify atomic structs, maps, or arrays, because of the risk of accidentally deleting values not included in the configuration.
  • JSONPatch: a CEL expression returning an array of JSONPatch{op, path, value} values. A jsonpatch.escapeKey() helper escapes / and ~ in keys, which you need for label keys like example.com/environment.

The docs’ sidecar example uses matchConditions for idempotence: it skips the mutation when an init container named mesh-proxy already exists, and sets reinvocationPolicy: IfNeeded so the policy can run again if later mutations change the object. A compact example in the same style adds a default owner label when one is missing:

apiVersion: admissionregistration.k8s.io/v1
kind: MutatingAdmissionPolicy
metadata:
  name: default-owner-label.policy.example.com
spec:
  failurePolicy: Fail
  reinvocationPolicy: IfNeeded
  matchConstraints:
    resourceRules:
    - apiGroups: ["apps"]
      apiVersions: ["v1"]
      operations: ["CREATE"]
      resources: ["deployments"]
  matchConditions:
  - name: owner-label-absent
    expression: "!has(object.metadata.labels) || !('owner' in object.metadata.labels)"
  mutations:
  - patchType: ApplyConfiguration
    applyConfiguration:
      expression: >
        Object{
          metadata: Object.metadata{
            labels: {"owner": "unassigned"}
          }
        }

This follows the documented schema, but I have not run it against a v1.36 or later cluster, so test it in a scratch namespace first. Pair it with a binding, exactly as with validation: a policy without a binding does nothing.

Pair mutation with validation

Do not trust a mutation to be the final word. A mutating policy can be overwritten by another mutator, and a client can always supply a value you did not expect. The robust design is a mutating policy to fill defaults and a ValidatingAdmissionPolicy to assert the final state, which is the same guidance the webhook documentation gives. The docs also note that for purely preventive rules, ValidatingAdmissionPolicy is the simpler and more effective alternative to a mutating policy.

Things that are exempt

Policies written through the REST API cannot intercept policy objects themselves. Per the docs, ValidatingAdmissionPolicy, its binding, and the mutating equivalents are exempt from policies created through the API, to prevent circular dependencies. The documentation also describes manifest-based admission control, where policies defined on disk can intercept those kinds, because a bad on-disk policy would not be unrecoverable the way a bad API-stored policy could be. Separately, the virtual authentication and authorization kinds such as TokenReview and SubjectAccessReview are exempt from all admission policies, including manifest-based ones, since intercepting them could lock the cluster out of its own auth path. The page adds that from v1.37 admission webhooks also exclude these virtual resources by default.

Failure Modes: Webhooks Versus In-Process Policy

A fair comparison starts with what actually goes wrong in production. The webhook failure catalogue is long, and the Kubernetes documentation itself supplies much of it.

Webhook failure modes

Availability coupling. With the default failurePolicy: Fail, an unreachable webhook rejects every matching request. The docs observe that this default can reject compliant requests during webhook downtime. Setting Ignore converts the outage into a policy hole instead, which is why the docs suggest letting mutating webhooks fail open and relying on a later validating step.

Self-referential deadlock. The good-practices page describes a webhook that requires a label that its own Deployment lacks. When the node running the webhook fails, rescheduling is blocked because the existing webhook rejects the replacement pods. The documented mitigation is to exclude the webhook’s own namespace with a namespaceSelector.

Dependency loops. Two webhooks that check each other’s pods can deadlock if both go down. A webhook that intercepts a networking or storage add-on it depends on can do the same. The recommendation includes using ValidatingAdmissionPolicies to avoid introducing dependencies.

Latency stacking. Every webhook call adds latency to API requests. Timeouts are bounded to 1 through 30 seconds with a default of 10, and a slow webhook at the default timeout can hold up every write it matches. During a control-plane incident, when clients retry aggressively, that amplifies load.

Operational surface. Certificates expire, CA bundles drift, rollouts of the webhook deployment collide with cluster upgrades, and RBAC must protect webhook configurations because they are powerful controllers.

What changes with in-process policy

The ValidatingAdmissionPolicy has no server, no certificate, no service, and no network path. A policy can only fail for reasons internal to the apiserver: a CEL error, a missing parameter, or a cost-budget overrun. All are deterministic, reproducible from the object and policy alone, and visible in status.typeChecking and in the audit trail. The availability of your policy equals the availability of your apiserver. In HA control planes, every apiserver replica evaluates the same stored policies independently.

That does not mean zero risk. A buggy policy with Fail and Deny can still block all deployments cluster-wide the moment it is saved, as no webhook outage is needed. The difference is that the blast radius is a configuration mistake you can revert with one command from a client that is exempt, instead of an infrastructure dependency you may not be able to restart. Keep break-glass access: know which credentials can delete a policy binding, and do not write a policy that blocks their own access path. The ValidatingAdmissionPolicyBinding kinds are exempt from API-created policies, so a bad policy cannot prevent you from deleting its binding.

Our chaos engineering guide for Kubernetes describes how to inject exactly the pod-kill and network faults that expose webhook fragility; running such an experiment against your existing webhooks is the quickest way to find out how many of your policies are single points of failure.

Trade-offs, Gotchas, and What Goes Wrong

ValidatingAdmissionPolicy is not a superset of webhooks. These are the limits that matter in practice.

No external data. CEL expressions see the request, the old object, the namespace object, a bound parameter, and the authorizer. They cannot query a vulnerability database, an image-signature service, a CMDB, or another cluster object beyond what a parameter resource encodes. Image signature verification and software bill-of-materials checks need a webhook-based tool. The rule of thumb: if the answer depends on something that is not in the request or a parameter, use a webhook.

Parameters are not a database. It is tempting to stuff long allow-lists into ConfigMaps and check image in params.data.... That works for small lists, but large parameter objects increase cost per evaluation, and the policy is evaluated per binding and parameter combination.

Reporting is thin. Kyverno-style policy reports and exception workflows are a product feature that the native API does not give you. You get audit annotations, warnings, and status.typeChecking. For compliance dashboards you will assemble tooling or adopt an engine on top.

Errors are not false. An absent field raises an error, not a clean negative. With failurePolicy: Fail that blocks the request; with Ignore it silently disables the check for that request, and the docs say failures defined by the failure policy are only enforced through the validation actions when the policy is Fail. If you choose Ignore for safety, you have also chosen blindness. Prefer Fail with careful has() handling, and alert on policy errors.

Overlapping bindings multiply work and confuse blame. If two bindings both match, the policy runs twice, and the denial message names the binding that failed. Keep selectors disjoint.

Updates hit legacy objects. As noted in the rollout discussion, a new constraint applies to UPDATE of existing resources, which can break operations such as scaling a non-compliant Deployment. Use warn and audit first, and write transition-aware rules.

Subtle CEL traps. has() checks field presence; the in operator checks map keys. Using the wrong one produces errors. Regex cost grows with string length. List literals carry a base cost of 40 units each time they appear in an expression, so hoist constants into variables. Equality on list-type set and map arrays ignores element order, which is convenient but surprising if you expect positional comparison.

Version drift. Everything here depends on your control plane version. The docs for parameterNotFoundAction currently show an example with v1alpha1 for the binding in one section while the rest of the page uses v1; treat docs examples as illustrative and rely on kubectl explain against your own cluster. Features such as manifest-based admission control and the v1.37 webhook default change are newer still, so verify them in the release notes for the version you run.

Admission is not runtime. Admission control validates objects at write time. It cannot see what a container does after it starts. Pair it with network-level controls; our guide to Kubernetes network policy egress covers the runtime side that admission cannot reach.

Choosing: Decision Framework and Comparison

Decision flowchart for choosing between ValidatingAdmissionPolicy, MutatingAdmissionPolicy and admission webhooks or policy engines

Figure 4: A decision path. External data or network calls point to webhooks or an engine, object mutation on 1.36 or newer points to MutatingAdmissionPolicy, and pure request logic points to ValidatingAdmissionPolicy.

Figure 4 encodes the argument. Start from what the rule needs, not from tooling preference. Rules that need data beyond the request leave the in-process path. Rules that change the object need a mutator, and the mutating policy route is the first choice when your control plane supports the stable feature. Everything else is a ValidatingAdmissionPolicy candidate, and should be rolled out through audit, warn, and deny.

The matrix below summarises the trade-offs. The ratings are my qualitative judgement from the documented mechanisms, not benchmark results.

Dimension ValidatingAdmissionPolicy Admission webhook (custom) Kyverno / Gatekeeper
Runs where Inside kube-apiserver Your service, over the network Engine pods, via webhook (with growing native CEL integration)
Availability dependency Apiserver only Webhook service, TLS, DNS Engine pods, TLS, DNS
Language CEL Anything Rego, YAML or CEL, depending on the engine
External data No Yes Yes, via engine features
Mutation Via MutatingAdmissionPolicy (stable v1.36 per docs) Yes Yes (Kyverno), limited elsewhere
Reporting and exceptions Audit annotations and warnings only You build it Built in
Ops overhead Low High Medium
Best for Structural guardrails, naming, labels, limits, tag rules Signatures, inventory lookups, bespoke logic Large rule libraries, reporting, multi-team governance

Two pragmatic combinations dominate real clusters. The first is layering: use native policy for the dozen universal guardrails that must never depend on a webhook being up, and an engine for the long tail needing reports, exceptions, or external data. The second is generation: Kyverno documents generating ValidatingAdmissionPolicies from its own policies, so you can keep its authoring model while moving enforcement in-process. Gatekeeper has also been adding a CEL-based path; check each project’s current documentation for exact support before committing. I have not verified the specifics of either integration in this article.

Practical Recommendations

Start by inventorying your webhooks. For each, ask whether the rule needs anything beyond the request and a small parameter. Webhook configurations whose servers only compare fields, check labels, or enforce limits are candidates for migration, and each migration deletes a service, a certificate, and an on-call page.

Write policies as small, single-purpose objects with descriptive names and a validation message or messageExpression that tells the engineer what to change. Put values in parameters, not in the logic, so one policy can serve many environments. Use variables to name anything you reference twice.

Always gate with matchConditions or selectors for the exemptions you truly intend, and exclude kube-system and your platform namespaces deliberately, rather than by accident. Test every policy with server-side dry runs, and read status.typeChecking after every apply.

Roll out in three steps per binding: Audit, Warn, then Deny, and watch the audit stream for evaluation errors as well as violations. Keep policy manifests in Git and apply them through your normal pipeline, with a CI job that runs the pass and fail fixtures.

The checklist before you flip a binding to Deny:

  • Type checking shows no warnings for the matched types.
  • Pass and fail manifests have been run through server-side dry run.
  • The parameter guard exists, and parameterNotFoundAction is set on purpose.
  • failurePolicy is Fail, and has() guards cover every optional field.
  • Existing objects were audited, and the team owning them has been warned.
  • A break-glass path to delete the binding is documented and tested.
  • Mutating logic, if any, is backed by a validating policy for the final state.

Frequently Asked Questions

What is a ValidatingAdmissionPolicy in Kubernetes?

A ValidatingAdmissionPolicy is an API object that defines admission rules in the Common Expression Language and runs them inside the kube-apiserver. It replaces many validating webhooks with a declarative, in-process check. A policy needs at least one ValidatingAdmissionPolicyBinding to take effect, and can take parameters from a ConfigMap or custom resource. It has been stable since Kubernetes v1.30.

Does ValidatingAdmissionPolicy replace OPA Gatekeeper and Kyverno?

Not entirely. It replaces webhooks for rules that depend only on the request, old object, namespace, and parameters. Gatekeeper and Kyverno still add reporting, exceptions, rule libraries, external data access, and, in Kyverno’s case, rich mutation and generation. Many teams run native policies for universal guardrails and an engine for the rest. Both engines also document or are building CEL-related integration, so check their current docs.

What happens if a CEL expression errors or a parameter is missing?

The policy’s failurePolicy decides. With Fail, the default, an evaluation error rejects the request. With Ignore, the error is dropped and the request proceeds. For missing parameters, the binding’s parameterNotFoundAction is required: Allow treats absence as a pass, while Deny applies the policy’s failure policy. Add a params != null validation to make missing parameters explicit.

How do Deny, Warn, and Audit differ?

validationActions on the binding choose enforcement. Deny rejects the request with the failure message. Warn admits it and returns the failure as an HTTP warning to the client. Audit admits it and records the failure in the audit event. You can combine Warn with Audit, or Deny with Audit, but not Deny with Warn, since that would report the same failure twice.

Is MutatingAdmissionPolicy stable?

According to the current Kubernetes documentation, MutatingAdmissionPolicy has been stable since v1.36 and is enabled by default, having first been available in v1.30. Mutations are written in CEL as either server-side-apply style apply configurations or JSON patches. If you run an older control plane, check your release notes and feature gates, as pre-stable behaviour and API versions differ.

How do I test a ValidatingAdmissionPolicy before enforcing it?

Create the policy and a binding with validationActions: [Warn] or [Audit], then use kubectl apply --dry-run=server with passing and failing manifests. Read status.typeChecking to catch undefined fields. For logic and regex, test expressions in the CEL Playground or a local CEL library. Promote the binding to Deny only after audit data shows no unexpected violations.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *