Falco 0.45 vs Tetragon vs Tracee: Kubernetes Runtime Threat Detection with eBPF
Every Kubernetes security team eventually asks the same question after an image scanner and an admission controller are in place: what happens when something malicious runs anyway? Static controls say nothing about a compromised dependency spawning a shell at 3 a.m., a container reading the node’s service account token, or a cryptominer reaching out to a pool. Closing that gap is the job of Falco runtime threat detection and its two main open source rivals, Tetragon and Tracee, all of which watch the kernel through eBPF.
The timing matters because the field just moved. Falco 0.45.0 shipped on 21 September 2026 with a new driver API, raw-byte rule matching and a hot-reload control endpoint, while Tetragon 1.7 added in-kernel CEL expressions and Tracee has had no stable release that I could verify since November 2025. Choosing between them is no longer a matter of picking the most popular logo.
This post gives you a working mental model of how each tool collects events, what its policy language can and cannot express, where enforcement is real versus aspirational, and how to route and tune alerts. You leave with a decision matrix and an opinionated recommendation.
What this covers: the eBPF data paths of all three tools, Falco 0.45 changes, rule and policy syntax side by side, detection versus enforcement, Falcosidekick routing, noise tuning, failure modes, and a pick-by-situation matrix.
Context and Background
Runtime security for containers rests on one observation: every container shares the host kernel, so every meaningful action a workload takes, from opening a file to connecting a socket, passes through a system call or a kernel function. If you can observe those transitions with enough context to know which pod, namespace and image they belong to, you can describe attacker behavior without modifying the application. That is the premise Falco launched on at Sysdig in 2016 and that the Cloud Native Computing Foundation (CNCF) now hosts as a graduated project.
The three tools reach the kernel differently. Falco historically intercepted raw system calls through a kernel module, then through eBPF probes, and now defaults to a “modern eBPF” probe that ships inside the Falco binary. Tetragon, built by Isovalent and developed within the Cilium project, attaches eBPF programs to kernel functions, tracepoints and Linux Security Module (LSM) hooks and evaluates filters inside the kernel. Tracee, maintained by Aqua Security, traces events with eBPF and processes them in a Go pipeline in userspace. If you want a deeper treatment of the Cilium side of the family, our earlier piece on Tetragon runtime security with eBPF walks through its policy engine, and the networking angle is covered in the Cilium eBPF service mesh, observability and security guide.
It helps to separate three jobs that vendors blur together. Detection means producing an alert when behavior matches a description. Enforcement means changing the outcome, by killing a process or denying a call. Forensics means preserving evidence, such as captured files or a process lineage, so a human can reconstruct an incident. Falco is overwhelmingly a detection tool. Tetragon can do all three, with enforcement as a first-class feature. Tracee is detection plus forensics, and its own documentation frames it as detection rather than blocking.
The other piece of context is the kernel itself. Container runtime hardening, such as seccomp profiles, AppArmor and user namespaces, shrinks what an attacker can do, and runtime detection tells you when they try anyway. Both layers matter, and the hardening side is discussed in our container runtime security hardening guide for containerd 2.3, CRI-O and Podman 6. Think of eBPF tooling as the smoke detector installed after the fire-resistant walls.
A word on sourcing. Version numbers, dates and feature lists below were checked against the upstream release pages and documentation on 6 October 2026. Where a vendor publishes no overhead benchmark, I say so rather than quoting a figure, because the published numbers for this category are mostly vendor-run on unspecified workloads.
Three Architectures for Watching the Kernel
The central difference between these tools is where the decision is made. Falco ships nearly every event to userspace and decides there, Tetragon decides inside the kernel and ships only what survived, and Tracee collects broadly and filters in its userspace pipeline according to policy. Each choice trades flexibility against cost and against how early a response can fire.

Figure 1: Three eBPF data paths. Falco streams a full event stream to userspace, Tetragon filters and acts in kernel, Tracee decodes and enriches in a userspace pipeline.
The diagram shows the three paths leaving the same kernel. In the Falco path, a driver copies system call events into per-CPU or shared ring buffers and a userspace library reconstructs state. In the Tetragon path, small programs attached to chosen hooks apply selectors before anything is sent. In the Tracee path, programs emit structured events that a Go engine decodes, enriches with container metadata and matches against policies and detections.
Falco: a full event stream with a rich userspace state machine
Falco is built on a set of libraries, libscap for capture and libsinsp for state and enrichment, collectively called libs. The driver feeds libscap, and libsinsp turns raw events into a model that knows processes, threads, file descriptors and, through the container plugin, container and Kubernetes metadata. The rule engine then evaluates boolean conditions over fields such as proc.name, fd.name, container.id and k8s.pod.name.
The consequence of this design is expressiveness at a cost. Because Falco sees the whole stream, a rule can correlate a file read with the process ancestry that led to it, using state the userspace model has built up. The price is that every event crosses the kernel boundary before any rule says yes or no, so buffer sizing and drop behavior become real operational concerns.
Falco supports two drivers today according to its documentation: the modern eBPF probe, which is the default, and a kernel module. The modern probe is embedded in the Falco binary, needs no separate build or download, and requires BPF ring buffer support and BTF (BPF Type Format) exposure; the docs describe kernel 5.8 as the typical minimum. The kernel module works on kernels from 3.10 but needs full privileges. The legacy eBPF probe, which compiled per-kernel objects, was removed in Falco 0.44.0 in May 2026 alongside the gVisor engine and the gRPC output.
Tetragon: policies compiled into in-kernel decisions
Tetragon is configured through a Kubernetes custom resource called a TracingPolicy. A policy names a hook point, a kernel function via kprobe, a tracepoint, a user-space function via uprobe, or an LSM hook, and attaches selectors that filter on arguments, process attributes, namespaces and binaries. The selectors are implemented in eBPF, which means non-matching events never leave the kernel.
The release history shows where the project is heading. Tetragon 1.7.0, released 29 April 2026, added environment variable retrieval, CEL (Common Expression Language) evaluation inside eBPF programs, a fentry sensor for function entry tracing, host-level selectors and parent binary matching. Version 1.7.1 followed on 25 August 2026 with 84 commits of stability work, including memory reporting fixes and CEL validation when loading policies in standalone mode. The in-kernel CEL support is the notable item because it narrows the expressiveness gap with userspace rule engines.
The documentation also carries a blunt warning that TracingPolicy is low-level and requires knowledge of the kernel and containers to avoid problems such as time-of-check to time-of-use (TOCTOU) races. That honesty is worth internalizing: you gain early, in-kernel decisions and you take on responsibility for choosing correct hook points.
Tracee: events first, detections second
Tracee began as a tracing tool and grew detection on top. Its documentation describes lifecycle, security, network, LSM and system call event categories, and states that its focus is detection and reporting rather than blocking. Policies are YAML objects with a name, a scope that says which workloads they cover, and rules that name events and optional data filters, for example restricting a file event to /tmp/*. The docs note a limit of 64 simultaneous policies and that each event type may appear only once within a policy.
Beyond policies, Tracee supports custom detection logic written in Go, and its artifact capture features, which can preserve files written or executed, are what set it apart for forensics. The caveat is momentum. The latest stable release on the project’s GitHub page is v0.24.1, published 19 November 2025, which puts the gap between that release and today at roughly eleven months. I could not verify any 2026 stable release, so treat the project’s maintenance trajectory as a risk to confirm before standardizing on it.
Falco 0.45 in Detail: What Changed and Why It Matters
Falco 0.45.0 is not a feature-heavy release, but it contains three changes that affect operators directly: a driver API bump, a change to how rule conditions treat bytes, and a hot-reload control surface. Reading the notes carefully before upgrading will save you a failed rollout on a node pool with mixed driver versions.
Component versions and the driver compatibility rule
The release page for 0.45.0 lists Falco libs 0.26.0, driver 11.0.0+driver, falcoctl 0.14.2, container plugin 0.7.4, falco-rules 5.2.0 and Helm chart 9.2.0. The driver API moved to a new major version, and the release states that older drivers are incompatible. In practice this matters most if you pin a kernel module build or distribute drivers through falcoctl: the driver and userspace must move together, and an agent upgrade that races a driver upgrade will leave nodes without detection.
Falco 0.44.0, released 26 May 2026, had already done the same thing to drivers, going to 10.2.0 and breaking compatibility with 0.43. Two consecutive major driver bumps in four months is a pattern, and it argues for treating Falco upgrades as a coordinated, canaried change rather than a routine image bump. A third notice in the 0.45 notes is that the memory used by auxiliary eBPF maps doubles per CPU, which the maintainers describe as being for eBPF safety. On large many-core nodes, check the memory request for the Falco DaemonSet after upgrading.
Raw byte matching and the UTF-8 change
The most important behavioral change is in the rule engine. Before 0.45, conditions were evaluated on text that had been through UTF-8 handling, so invalid byte sequences were replaced with the Unicode replacement character. Falco 0.45 evaluates conditions on raw field bytes and adds \xHH escape sequences so a rule can match invalid UTF-8 and non-printable characters, while output formatting escapes control characters so logs stay readable.
This is a meaningful detection improvement. Attackers use odd encodings and control characters in command lines and file names to evade string matches or to confuse log pipelines. It is also a compatibility risk the release notes flag explicitly: an existing rule that matched the replacement character may stop behaving as before, and downstream consumers must handle the new escaped output. Run your rule set through validation against 0.45 in a staging cluster and diff alert volume before you promote.
Hot reload, packaging and hardening fixes
Falco 0.45.0 adds incubating endpoints for reload management: a GET /reload for status and an optional POST /reload exposed through a protected Unix socket on Linux. The motivation is operational. Rule updates delivered by falcoctl or a GitOps pipeline previously needed a signal or restart with little feedback on whether the new rule set loaded. A status endpoint lets automation confirm the outcome.
The release also fixed several security-relevant issues that deserve a mention because a detection agent runs with high privilege on every node. Falco now uses the absolute modprobe path instead of a PATH lookup when falling back to loading the kernel module, closing an untrusted search path weakness classified as CWE-426. It rejects symlinks when creating the pidfile and fixes a use-after-free. Package upgrades for DEB and RPM now restore the selected services and drivers, which fixes upgrades that left a host in a non-functional state. The obsolete container_engines.* configuration was removed, so check your Helm values for leftovers.
What 0.44 set up
To understand 0.45 you need 0.44. That release added the oneof, anyof and allof list modifiers for comparison operators, which let a rule test one field against many values without repeating the operator. It added a hard cap on capture file size, capture.max_file_size_mb, accepting 0 to 1,048,576 MB, made the rule loader reject unknown top-level keys so typos fail loudly, and rewrote /proc parsers plus added kernel-side BPF iterators to speed up process tree lookups. A 0.44.1 patch on 11 June 2026 added a way to disable BPF iterators and fixed several issues related to them. If you run on an older or vendor-patched kernel and saw odd startup behavior, that is the knob to know about.
Rule Languages Side by Side: Writing the Same Detection Three Ways
The clearest way to compare the tools is to express one detection in each syntax. Take a common requirement: alert when an interactive shell starts inside a container. The example below is illustrative and simplified, not copied from any shipped rule set, but the field names and structure follow each project’s documented syntax.
Falco: boolean conditions, macros and lists
A Falco rule is a YAML object with a name, description, condition, output and priority. Conditions are boolean predicates over event fields, and macros and lists provide reuse. Priorities run from EMERGENCY down to DEBUG.
- list: shell_binaries
items: [bash, sh, zsh, dash, ksh, csh, tcsh]
- macro: spawned_process
condition: (evt.type in (execve, execveat) and evt.dir = <)
- rule: Terminal shell in container
desc: A shell was spawned in a container with an attached terminal
condition: >
spawned_process and container.id != host
and proc.name in (shell_binaries) and proc.tty != 0
output: >
Shell in container (user=%user.name pod=%k8s.pod.name
ns=%k8s.ns.name image=%container.image.repository
cmd=%proc.cmdline parent=%proc.pname)
priority: NOTICE
tags: [container, shell, mitre_execution]
Three things stand out. The condition is a single expression, so complexity grows inside the string. Exceptions are a separate rule field, which is how you carve out known-good behavior without editing the base rule. And the output is a template, which means the quality of your alert depends on the fields you chose to interpolate. With 0.44 and later, the oneof, anyof and allof modifiers shorten rules that compare a field to multiple values.
The upside of this model is a large community rule set. The falco-rules package, at version 5.2.0 alongside Falco 0.45.0, packages curated rules in maturity tiers, and falcoctl can fetch and update them. You start with coverage on day one and tune from there.
Tetragon: a TracingPolicy that names a hook and selectors
Tetragon expresses the same intent as a hook plus selectors. Because it works at kernel function level, you name the function that matters, here the exec path, and filter on the binary. The snippet is illustrative and follows the documented TracingPolicy shape; confirm field names against the version you deploy.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: shell-exec-in-prod
spec:
kprobes:
- call: "security_bprm_check"
syscall: false
args:
- index: 0
type: "linux_binprm"
selectors:
- matchNamespaces:
- namespace: Pid
operator: NotIn
values:
- "host_ns"
matchBinaries:
- operator: "In"
values:
- "/bin/bash"
- "/bin/sh"
matchActions:
- action: Post
The structure differs from Falco in a telling way. There is no output template, because Tetragon emits structured events, with process and Kubernetes context attached, through its gRPC stream and JSON logs. And there is a matchActions block, which is where detection turns into enforcement by changing Post to something stronger. Selectors are evaluated by eBPF in the kernel, so the policy author is also the performance engineer.
Tracee: policy scope plus events and filters
Tracee’s policy is simpler by design. It names the events to trace and filters on event data. The fragment below uses the documented v1beta1 policy shape with a built-in event; treat the specific event and filter as illustrative.
apiVersion: tracee.aquasec.com/v1beta1
kind: Policy
metadata:
name: exec-from-tmp
annotations:
description: Executions of binaries from temporary directories
spec:
scope:
- container
rules:
- event: sched_process_exec
filters:
- data.pathname=/tmp/*
Where Falco and Tetragon encode a verdict, a Tracee policy mostly encodes a collection request, and verdict logic lives in detections or in whatever consumes the event stream. That is a defensible separation, and it suits teams that already have a pipeline and want clean, richly typed events. It is less convenient if you want a rule file that goes straight to an on-call page.
What the syntax tells you about intent
The three snippets reveal three different products. Falco is a rule engine with a content ecosystem. Tetragon is a policy compiler for the kernel with optional enforcement. Tracee is an event source with policy-based collection. When teams compare them on feature checklists they miss this, and end up asking one tool to be another.
Detection Versus Enforcement: Where Blocking Is Real
Detection and enforcement are separated by time. A detector learns about an action after the kernel has copied an event out, so the best it can do is raise an alert and let something else respond. An enforcer sits on the decision path and can change the result. The distinction determines what a compromised pod can accomplish before you react.

Figure 2: Detection path versus enforcement path for the same suspicious exec. Detection produces an alert after the event is copied out; in-kernel enforcement acts on the call itself.
The sequence diagram contrasts the paths. On the detection side, the kernel hook copies an event to a buffer, userspace evaluates it, an alert fires and a responder acts later, a delay that can be seconds. On the enforcement side, a selector matches inside the kernel and the response is applied to the call.
How Tetragon enforcement works, and its documented limits
Tetragon’s documentation names two mechanisms. The first overrides a function’s return value, so the guarded function never runs and the caller receives an error. This works for system calls and for security check functions. The second sends a signal, typically SIGKILL, to the process that matched.
The documentation is candid about a limit of the signal approach: sending SIGKILL does not always stop the operation the process was performing. A SIGKILL delivered during a write() does not guarantee the data will not reach disk, though it does guarantee the process is terminated. The guidance is to combine the Signal action with the Override action to ensure the operation is not completed. That is the correct mental model: kill for containment, override for prevention, both for a hard block.
Enforcement also raises the cost of mistakes. A detection rule with a false positive produces a noisy alert. An enforcement policy with a false positive terminates a production process. For that reason, start every enforcement policy in observe mode, with Post actions only, and promote it after weeks of clean data.
Falco’s position: detect, then hand off
Falco does not block. Its answer to response is an ecosystem: alerts leave the agent, pass through Falcosidekick, and reach a responder such as Falco Talon, a response engine listed among Falcosidekick’s outputs, or any serverless function or workflow system. This keeps the privileged, high-volume agent simple and auditable, and it keeps response logic in a place where it can be reviewed, rate-limited and tested.
The latency difference is a design consequence, not a bug. If your threat model includes fast, scripted attacks where seconds matter, such as a payload that encrypts files on first execution, a detect-and-respond loop can lose the race. If your threat model is human-operated intrusion, where an attacker spends minutes on discovery, alerting within a second is usually early enough.
Tracee’s position: observe and preserve
Tracee’s stated focus is detection and reporting, and its distinctive feature is evidence capture. For an incident responder, a copy of the dropped binary or the written file is often worth more than the alert that pointed to it. If forensics is a hard requirement, this is the area where Tracee deserves a serious look, subject to the maintenance caveat noted earlier.
Detection versus enforcement at a glance
| Capability | Falco 0.45 | Tetragon 1.7 | Tracee 0.24 |
|---|---|---|---|
| Primary role | Detection | Detection and enforcement | Detection and forensics |
| Where filtering happens | Userspace rule engine | In-kernel selectors | Userspace policies and detections |
| Block or kill | No, via external responder | SIGKILL and return override | Not its stated focus |
| Policy format | YAML rules with conditions | Kubernetes TracingPolicy CRD | YAML policy, Go detections |
| Alert routing | Falcosidekick, many outputs | gRPC, JSON, export integrations | Multiple output formats |
| Default kernel path | Modern eBPF, or kmod | eBPF at kprobe, LSM, fentry | eBPF |
| Latest stable checked | 0.45.0, 21 Sep 2026 | 1.7.1, 25 Aug 2026 | 0.24.1, 19 Nov 2025 |
The table compares documented capabilities, not quality. A row where a tool says “no” is not a defect if that tool’s design deliberately excludes the capability.
From Kernel Event to Pager: Falco’s Alert Pipeline
Alert routing is where most operational pain lives, and it is Falco’s strongest area. The path from event to responder has more stages than the rule syntax suggests, and each stage is a place to lose, enrich or suppress signal.

Figure 3: Falco’s pipeline. The driver feeds libscap, libsinsp adds state and container metadata, the rule engine emits alerts, and Falcosidekick fans them out to chat, logging, SIEM and response targets.
The figure shows the userspace stages. The container plugin contributes container and Kubernetes metadata into libsinsp, which is why a rule can reference a pod name even though the kernel only knows process identifiers. Alerts emitted by the engine can go to stdout, files or an HTTP endpoint, and the HTTP output is what typically connects to Falcosidekick.
Falcosidekick as a fleet-wide forwarder
Falcosidekick describes itself as a proxy forwarder acting as a central point for any fleet of Falco instances. It can add custom fields to alerts, filter by severity, and expose its own metrics. The destinations documented include chat platforms such as Slack, Discord, Teams and Mattermost, monitoring systems such as Datadog and Prometheus, alerting systems such as PagerDuty, Opsgenie and Alertmanager, logging backends such as Elasticsearch, Loki, CloudWatch and syslog, object storage on AWS and GCP, messaging systems such as Kafka, RabbitMQ and SQS, serverless targets, AWS Security Lake and Falco Talon.
The design choice worth copying is per-output minimum priority. Send everything at NOTICE and above to a log store for hunting, WARNING and above to chat, and CRITICAL and above to the pager. Those thresholds are a cheap, effective noise control that does not require touching a single rule.
Falcosidekick UI and Falco Talon
Falcosidekick UI provides a dashboard with event statistics, which is useful in the first weeks of tuning because it shows which rules dominate volume. Falco Talon sits at the other end, taking alerts and running actions, such as labeling or terminating a pod. Because response is an external, auditable component, you can gate it behind approvals and dry-run it, a property that in-kernel enforcement does not easily offer.
Tuning Noise, Measuring Overhead and Operating the Agents
A runtime detector that nobody trusts is worse than none, because the alerts train people to ignore the channel. Operating these tools is mostly a noise-management and capacity problem, and the way to approach it is the same across all three.
A tuning loop that works for Falco
Start with the default rules in alert-only mode for two weeks, route everything to a log store and nothing to a pager. Rank rules by volume using Falcosidekick UI or your log queries. For each of the top five noisy rules, decide whether the behavior is legitimate for that workload, and if it is, add an exception scoped to the narrowest fields that identify it, usually image repository plus process name plus command line, never just the process name.
Scope matters because an exception on proc.name = curl silences the one rule and also hides the attacker who uses curl. Prefer exceptions keyed on the workload identity, namespace and image, and review them on a calendar. Treat the exception list as code in version control with a reviewer, because it is the place where detection coverage quietly erodes.
Next, enrich outputs. A good Falco output contains the pod, namespace, image, user, the full command line, and the parent process, because the responder’s first three questions are always what, where and who started it. Add custom fields in Falcosidekick for cluster name and environment so that multi-cluster alerts are attributable without a lookup.
Tuning Tetragon and Tracee
Tetragon noise is controlled at the policy, not the rule, so the lever is the selector. Scope policies by namespace, pod label or binary, and prefer narrower hooks. A kprobe on a hot kernel function with a loose selector will emit enormous volumes, and because the work occurs in the kernel path, a poorly scoped policy costs CPU for every call, not only for matching ones. This is the reason the project’s own warning about low-level policy authoring is worth repeating to reviewers.
Tracee’s lever is scope and event selection. Restrict policies to the workloads that need them, remember the limit of 64 simultaneous policies and one definition per event type per policy, and use data filters on paths to cut volume. Because Tracee leans on a userspace pipeline, the cost of an over-broad policy lands in the agent’s CPU and memory instead of the kernel path.
Overhead: what is sourced and what is not
Readers always ask what the CPU overhead is, and the honest answer is that it depends on event rate, which is a property of your workload, not the tool. The Falco documentation says only that the modern eBPF probe offers better performance and maintainability than the kernel module and lets you tune buffer layout, including a single shared buffer instead of one per CPU, through engine.modern_ebpf.buf_size_preset. I did not find a vendor-published, methodology-complete overhead benchmark for any of the three tools at the versions discussed, so this post quotes no overhead percentages.
What you can do is measure your own. Run a representative load, for example an HTTP service with a known request rate, on two identical nodes, with and without the agent, and compare p99 latency, CPU and memory. Then repeat with a syscall-heavy job such as a build or a database checkpoint, because the agents behave differently when the event rate is ten times higher. The numbers below are an illustrative template for how to record results, not measurements.
| Scenario | Metric | Baseline | With agent | Delta |
|---|---|---|---|---|
| Web service at fixed RPS | p99 latency | measure | measure | compute |
| Syscall-heavy job | wall clock | measure | measure | compute |
| Idle node | agent RSS and CPU | n/a | measure | n/a |
| Burst of exec events | dropped events | n/a | measure | n/a |
The last row matters most for Falco. When the event rate exceeds what userspace can drain, the driver drops events, and a dropped event is a blind spot. Falco has a documented drop-handling mechanism and metrics, so alert on non-zero drops, since an attacker who can generate noise can try to hide in it. The same discipline applies to the other tools: any tool that sheds load under pressure needs a metric you watch.
Privileges, kernel features and node lifecycle
All three agents run with elevated privileges because loading eBPF programs requires them. For Falco’s modern eBPF probe, the documentation lists the least-privileged capabilities: CAP_SYS_RESOURCE for memory locking, CAP_SYS_PTRACE for process environment, and CAP_BPF and CAP_PERFMON, which may replace the broader CAP_SYS_ADMIN on kernels 5.8 and later. The kernel module path needs full privileges. If your platform team enforces Pod Security Standards, this is the argument for the modern probe even before performance enters the conversation.
Kernel compatibility is the second operational variable. The modern probe needs BPF ring buffers and BTF, so older enterprise distributions may not qualify, while the kernel module reaches back to 3.10. Managed Kubernetes services generally ship recent kernels, but custom node images and edge fleets are where you find surprises. Verify with a canary node pool before a fleet rollout.
Upgrades as a first-class risk
Given two consecutive Falco driver API bumps and the 0.45 change in rule byte handling, build an upgrade runbook: canary pool, run old and new in parallel where possible, diff alert counts per rule, then promote. For Tetragon, read the release notes for removed APIs, since recent releases removed legacy stack-trace APIs and deprecated certain gRPC methods. A security agent that silently stops working after an upgrade is a failure you will discover during an incident.
Choosing Between Them: A Decision Framework
No tool wins on every axis, and the right answer is often two of them. The decision depends on three questions: do you need to block in the kernel, do you value a mature rule ecosystem and broad alert routing, and do you need file-level evidence for investigations.

Figure 4: A decision flow. Kernel-level blocking points to Tetragon, ecosystem breadth points to Falco, and evidence capture points to Tracee, with Falco plus Tetragon as a common pairing.
The flow encodes an opinion: if enforcement is a hard requirement, start with Tetragon, and if it is not, Falco’s rules and routing are the default. Tracee earns its place when forensic artifacts are the priority and you accept the maintenance risk.
Decision matrix by use case
| Use case | Best fit | Why | Watch out for |
|---|---|---|---|
| Fleet-wide detection with curated rules on day one | Falco | Rule set, tiers, falcoctl updates, Sidekick routing | Driver API churn, rule tuning effort |
| Block known-bad behavior such as a shell in prod | Tetragon | SIGKILL plus return override in kernel | Policy mistakes kill workloads |
| Capture dropped files and artifacts for IR | Tracee | Artifact capture and typed events | No verified 2026 stable release |
| Network-aware identity and policy with Cilium | Tetragon | Same ecosystem and metadata | Needs kernel-savvy authors |
| Compliance evidence of runtime monitoring | Falco | Familiar, widely audited, mature outputs | Alerts are not proof of prevention |
| Edge or old kernels below 5.8 | Falco kmod | Works from kernel 3.10 | Needs full privileges |
Pairing Falco and Tetragon
A defensible architecture uses Falco as the broad detection layer, with its rules and routing, and Tetragon as a narrow enforcement layer for a short list of high-confidence behaviors: spawning a shell in a production namespace, writing to specific protected paths, loading kernel modules from a container. The two overlap on events, which is fine, because the cost is mostly CPU on nodes and some duplicate alerts you can deduplicate in Falcosidekick or the SIEM.
The pairing also matches how mature teams staff the work. The detection layer is owned by a security operations team that tunes rules. The enforcement layer is owned jointly with platform engineering, because a bad policy is an outage. Giving each layer a different owner and change process reduces the chance that an experimental rule kills a payment service.
Where eBPF security meets the AI agent era
One emerging reason to care about runtime signals is AI agents that execute tools and code inside containers. A prompt-injected agent that is tricked into running a shell or exfiltrating a token produces exactly the kernel-level behavior these tools describe, regardless of what the model was told. Our write-up on agentic AI security and prompt injection covers the model-side controls, and runtime detection is the backstop when those controls fail. Sandboxing an agent’s execution environment and alerting on unexpected process or network behavior is a natural fit for any of the three tools.
Trade-offs, Gotchas, and What Goes Wrong
Every one of these tools fails in characteristic ways, and most of the failures are operational rather than technical. Naming them in advance is cheaper than discovering them during an incident.
Dropped events create silent blind spots. Falco’s userspace model depends on draining buffers fast enough. Under bursty load, events are dropped, and an attacker who can spawn many cheap syscalls can try to hide behind the noise. Alert on drop counters and size buffers deliberately.
The kernel is a moving target. eBPF programs depend on kernel internals, BTF and hook availability. A node image upgrade can change function signatures or disable features, so each agent needs a compatibility test on every new kernel your fleet adopts. This is also where driver API majors bite: userspace and driver must match.
Time-of-check to time-of-use gaps persist. Hooking a system call at entry and inspecting user-space memory can be raced by another thread that changes the argument after the check. Tetragon’s documentation names TOCTOU explicitly. Hooking at LSM points, where arguments are resolved kernel objects, reduces the risk, and that is a reason to prefer LSM hooks for enforcement.
Container escapes change the question. If an attacker reaches the host with root, they may be able to disable or tamper with the agent. Detection agents are not a substitute for node hardening, and enforcement does not help against a kernel exploit that runs below the hook. Pair these tools with the hardening controls described in our container runtime hardening guide.
Rules rot. Detection content written against yesterday’s attacker tooling misses today’s, and exceptions accumulate until a rule no longer fires on anything. Schedule quarterly reviews, and test rules with a known-bad simulation after each change.
Alert fatigue defeats the investment. A pager that fires forty times a day gets muted. Use severity-based routing, aggregate duplicates, and treat the percentage of alerts that lead to action as a metric you track.
Enforcement without observation is reckless. Never ship a blocking policy without a long observation period. A SIGKILL on a legitimate process during a deploy looks like a platform outage, and the postmortem will name your security team.
Single-vendor risk. Tracee’s release cadence and Tetragon’s coupling to the Cilium project are both things to weigh. Falco’s graduated CNCF status and multi-vendor contribution base reduce that risk, though not to zero.
Practical Recommendations
For most Kubernetes platforms in late 2026, I recommend starting with Falco 0.45 as the detection baseline, the modern eBPF driver in least-privileged mode, Falcosidekick for routing, and a canaried upgrade process. It has the widest rule coverage, the clearest alerting path and a mature community. Add Tetragon when you have a short list of behaviors you are prepared to block, and add Tracee only after confirming its maintenance status against your own risk tolerance.
Sequence the work so that each step produces value on its own. Week one is install and observe. Weeks two to four are tuning, with exceptions scoped to workload identity. Month two is routing, with severity thresholds per destination and one owned on-call path. Month three is the first enforcement policy, in observe mode, promoted only after clean data.
Checklist before calling the rollout done:
- Confirm node kernels support the modern eBPF probe, or document the kernel module exception.
- Pin Falco, libs, driver and falco-rules versions together and canary every upgrade.
- Run the 0.45 rule validation in staging and diff alert volume against 0.44, given the raw byte matching change.
- Alert on dropped events and on agent health, not only on detections.
- Route by severity through Falcosidekick and measure the actioned-alert ratio.
- Keep exceptions in version control, scoped to image and namespace, with a review date.
- Start every enforcement policy with
Postactions and promote with a reviewer. - Measure your own overhead with the template above before fleet-wide rollout.
- Rehearse one incident end to end: simulate a shell in a container, trace the alert to the responder.
Frequently Asked Questions
Is Falco or Tetragon better for Kubernetes runtime security?
Neither is better in general. Falco is stronger for detection breadth, with a curated rule set, rich alert routing through Falcosidekick and a large community. Tetragon is stronger when you need to block behavior in the kernel, using SIGKILL and return-value override from a TracingPolicy. Many teams run Falco for detection and add Tetragon for a short list of high-confidence enforcement policies, accepting some overlap in the events both observe.
What does Falco 0.45 change compared with 0.44?
Falco 0.45.0, released 21 September 2026, moves the driver to 11.0.0 and libs to 0.26.0, evaluates rule conditions on raw field bytes with new \xHH escapes, adds incubating /reload endpoints for hot reload management, doubles auxiliary eBPF map memory per CPU, and removes the obsolete container_engines.* configuration. Older drivers are incompatible, so upgrade the driver and agent together and revalidate rules that matched replacement characters.
Does Falco block malicious processes?
No. Falco is a detection engine and does not terminate processes or deny system calls. Its model is to emit alerts, forward them with Falcosidekick, and let a response component such as Falco Talon, a serverless function or a workflow tool take action. That approach keeps the privileged agent simple, though a response loop takes longer than in-kernel enforcement, which matters for fast automated attacks.
How much overhead do eBPF security agents add?
It depends on your event rate, and I found no methodology-complete vendor benchmark at these versions, so I quote no percentage. Falco documents that its modern eBPF probe improves performance over the kernel module and allows buffer tuning. Measure your own: run the same workload with and without the agent, compare p99 latency and CPU, and test a syscall-heavy job, watching drop counters throughout.
Which kernel version do I need for Falco’s modern eBPF probe?
The Falco documentation gives 5.8 as the typical minimum for both x86_64 and aarch64, with the requirement being BPF ring buffer support and BTF exposure, so some kernels with backports below 5.8 can qualify. The kernel module works from kernel 3.10 but needs full privileges. If your nodes run older enterprise kernels, validate on a canary pool before choosing a driver.
Is Tracee still maintained in 2026?
I could verify only that the latest stable release on its GitHub releases page is v0.24.1, published 19 November 2025, with v0.24.0 five days earlier. The repository may have unreleased activity I could not confirm. Because that is roughly eleven months without a verified stable release, evaluate the project’s current commit activity and your tolerance for maintenance risk before standardizing on it.
Further Reading
- Cilium and Tetragon runtime security with eBPF for a deeper treatment of TracingPolicy and enforcement.
- Cilium eBPF service mesh, networking, observability and security for the networking half of the eBPF stack.
- Container runtime security hardening with containerd 2.3, CRI-O and Podman 6 for the preventive layer beneath runtime detection.
- Agentic AI security and prompt injection for why runtime signals matter for AI agent workloads.
- Falco 0.45.0 release notes on GitHub and the Falco documentation for primary sources.
- Tetragon documentation and the Tracee documentation for the other two projects.
By Riju — about
