Kepler 0.12: Kubernetes Pod Energy and Carbon Metrics Including GPU MIG Power Attribution

Kepler 0.12: Kubernetes Pod Energy and Carbon Metrics Including GPU MIG Power Attribution

Kepler 0.12: Kubernetes Pod Energy and Carbon Metrics Including GPU MIG Power Attribution

Nobody can meter a pod. A server has power sensors at the socket, the DIMM bank, the GPU board and sometimes the chassis, but no sensor knows what a container is. Every figure you see for Kepler Kubernetes energy per pod is therefore an attribution: a measured node-level quantity divided among workloads by a model, and the quality of that model decides whether your GreenOps dashboard is evidence or decoration.

Kepler 0.12.0 matters because it extends that attribution to the hardest case, shared GPUs. According to the release notes it adds MIG (Multi-Instance GPU) power attribution through dcgm-exporter, multi-architecture ARM64 images, and an hwmon fallback for hosts without RAPL. This post explains what Kepler actually does today, which parts of older descriptions (eBPF, trained ML models) no longer apply, how to query the metrics, how to turn joules into carbon, and where the numbers should not be trusted. You will leave with a mental model of the pipeline, working PromQL, and a checklist for deploying it with the right caveats.

What this covers: the post-rewrite architecture, the attribution formulas, GPU and MIG handling, metric names and PromQL, the carbon step Kepler does not do for you, and the accuracy limits.

Context and Background

Kepler (Kubernetes-based Efficient Power Level Exporter) is a CNCF sandbox project, accepted on 17 May 2023, and described by the CNCF as a Prometheus exporter that measures energy consumption at the pod and node level by monitoring hardware sensors and attributing usage to containers. It runs as a DaemonSet, so one agent per node, and publishes its series on a metrics endpoint that Prometheus scrapes. If you have read about Kepler before 2025, you read about a different program.

The original line (up to v0.9.x) used eBPF probes to count CPU cycles, instructions and cache misses per process, then fed those counters into trained regression models when real power sensors were absent. The project README now states that the 0.10.0+ line is a complete rewrite, that legacy 0.9.x is frozen with no further support, and that new users should adopt 0.10.0 or later. The rewrite is written in Go, detects RAPL zones dynamically instead of hardcoding them, and attributes power from active CPU usage. It also drops the elevated capabilities, CAP_SYSADMIN and CAP_BPF, that the eBPF design required, and needs only read-only access to host /proc and /sys according to the README.

That history explains a confusing amount of online material. A 2024 CNCF TAG Environmental Sustainability post on idle power and Kepler metrics for public cloud describes a trained dynamic power model for VMs that lack RAPL, citing idle power at 20 to 60 percent of maximum-utilization power. That is the legacy design. The current docs describe a simpler, sensor-first pipeline. When a tutorial mentions BPF programs, a model server or kepler_container_joules_total, check its date before copying anything.

The motivation has not changed. Cloud bills expose dollars, not joules, and the FinOps and GreenOps approach to cost and carbon-aware scheduling needs a per-workload energy signal to be more than a slogan. Kepler is the most visible open-source attempt to supply that signal inside Kubernetes. The same measurement problem sits behind public estimates of AI inference cost, which the post on AI query energy and water footprint with sources traced examines from the other end, starting from published figures rather than from a node.

A word on the release itself. The GitHub release page for v0.12.0 lists MIG power attribution via dcgm-exporter, energy readings for the hwmon power meter, multi-architecture ARM64 builds and images, HWMON as a backup when RAPL is unavailable, a configurable CPU meter selection (cpu.preferredMeters), VM process detection for VirtualBox, VMware and Xen, and liveness and readiness probes, with 20 contributors of whom 13 were first-time. I could not independently confirm the publication date: my fetch of the page returned a date that is inconsistent with the 0.10 to 0.11 sequence, so this post does not state one. The repository README I read still showed v0.11.4 in an example, so check the tag you actually deploy.

Kepler Kubernetes Energy: The Reference Pipeline

Kepler reads cumulative energy counters from hardware, reads per-process CPU time from procfs, splits the node’s energy into active and idle parts by CPU utilization, and distributes the active part to processes, containers, pods and VMs in proportion to CPU time. GPU power is added as a separate meter. Everything is published as Prometheus counters (joules) and gauges (watts).

Kepler Kubernetes energy pipeline from power meters and procfs through the monitor loop to Prometheus

Figure 1: Kepler Kubernetes energy pipeline. Hardware meters and procfs feed a periodic monitor loop; attribution results leave through the Prometheus exporter.

Figure 1 shows the data path. Power meters and the resource informer are independent inputs. The monitor loop, which the configuration reference sets to a 5 second interval by default, snapshots both, computes deltas since the previous snapshot, runs the attribution and updates the exported series. A staleness window of 500 ms (default) lets the exporter reuse a recent computation rather than recompute on every scrape. The informer combines process data from /proc with pod metadata, which Kepler obtains by polling the kubelet every 15 seconds by default or by using the API server.

Power meters: RAPL first, hwmon second

The CPU meter is chosen by an ordered list. The default is cpu.preferredMeters: ["rapl", "hwmon"], and Kepler walks the list, initializing each backend until one reports zones. A third value, fake, produces synthetic readings for development and must never feed a real dashboard.

RAPL (Running Average Power Limit) is an Intel mechanism, exposed on Linux through the powercap framework in sysfs. Per the kernel documentation, each zone has an energy_uj file holding a cumulative energy counter in microjoules and a max_energy_range_uj file giving the range of that counter, which means the counter wraps and any reader must handle rollover. The kernel documentation also notes that RAPL provides no instantaneous power value, so there is no power attribute: power is always derived as the energy delta divided by elapsed time. Zones map to the package, with subzones for core and uncore parts, and on many servers a DRAM zone. In Kepler you can restrict this with rapl.zones, whose empty default enables everything.

When RAPL is absent, the hwmon backend reads Linux hardware-monitoring sensors. The configuration reference places it under the experimental.hwmon block, with zones and chipRules that pair voltage and current channels (for example a chip such as an INA3221, where useSameIndex pairs in{N} with curr{N}). This is aimed at boards where power is derived from a voltage and a current reading, which is how many ARM and embedded systems expose it. It is a useful addition for the ARM64 builds that arrive in the same release, but it is an experimental path and worth validating against a wall meter before trusting it.

An optional third source is platform power from a BMC via Redfish, which the docs list as experimental and disabled by default. It exposes kepler_platform_watts with BMC and chassis labels, and it is the only path that sees fans, PSU losses and storage that RAPL cannot. It requires a configuration file mapping node names to BMC endpoints and credentials, so treat that file as a secret.

The attribution formula

The developer guide states one equation for every workload type: workload power equals the workload’s CPU time delta divided by the node’s CPU time delta, multiplied by the node’s active power. The split between active and idle uses the node CPU usage ratio. Active energy is the total energy delta times the usage ratio, and idle energy is the remainder.

The guide’s own worked example: hardware reports 40 W and the node is 25 percent utilized, so active power is 10 W and idle power is 30 W. A process that used 100 ms of a 1000 ms node CPU time budget, with 20 W of active power, receives 2 W. Containers sum their processes, pods sum their containers, and a VM’s CPU time is that of its hypervisor process.

Kepler Kubernetes energy attribution: node energy split into active and idle, then distributed to process, container and pod

Figure 2: Kepler attribution. Only active energy is distributed to workloads; idle energy stays at node level.

This is the single most important design choice, and it is easy to miss. Idle energy is not charged to pods. Summing kepler_pod_cpu_watts across a node will therefore always fall short of kepler_node_cpu_watts, by exactly the idle component. If your carbon report sums pods and compares to the facility meter, the gap is not an error, it is a policy decision you need to make explicit. Section “Practical Recommendations” returns to this.

How Kepler Kubernetes Energy Attribution Works in Detail

The pipeline above hides several mechanisms that determine accuracy. This section walks through the code paths that matter: counter handling, the usage-ratio split, the process-to-pod mapping, and terminated workloads.

Counters, deltas and the usage ratio

Every exported joule series is a Prometheus counter and every watt series a gauge computed from deltas. For the node, the metric families include kepler_node_cpu_joules_total, kepler_node_cpu_active_joules_total, kepler_node_cpu_idle_joules_total, and their _watts gauges, each labelled by zone and path. The usage ratio itself is exported as kepler_node_cpu_usage_ratio, a value between 0.0 and 1.0, which lets you reproduce Kepler’s split by hand and is the cheapest sanity check available.

Because the split is proportional to utilization, the model implies that a node’s power is linear in utilization between an idle floor and the measured value. Real servers are not linear. Frequency scaling, turbo states, SMT sharing and memory-bound versus compute-bound code all move watts per CPU-second. The model is correct in total at the node, since total energy comes from hardware, and approximate in the split. That distinction, exact in aggregate and heuristic in distribution, is the right way to read every Kepler number.

Process, container, pod and VM mapping

Processes are discovered from /proc and tagged with a container ID derived from the cgroup path, then joined to pod identity from the kubelet. The release notes for v0.11.3, per a release aggregator, mention a fix so the kubepods pattern is compatible with Guaranteed QoS pods, a reminder that cgroup path parsing is a real source of mislabelled or unattributed series. Process metrics carry pid, comm, exe, type, container_id and vm_id, which is high cardinality; the documentation lists them but you should not scrape them into long-term storage on large nodes without dropping labels.

Containers aggregate their processes and pods aggregate containers. The pod series are kepler_pod_cpu_joules_total and kepler_pod_cpu_watts, labelled by pod_id, pod_name, pod_namespace, state and zone. Note the state label, which distinguishes running from terminated workloads, and the zone, which means a pod’s CPU energy is reported per RAPL zone (package, core, DRAM and so on) rather than as one number. When you build a dashboard, decide which zones you sum. Package already includes core, so summing every zone double counts.

Terminated workloads and short-lived processes

Two settings control what is lost. Processes that live shorter than the monitor interval are ignored, per the configuration reference. Terminated workloads are retained up to monitor.maxTerminated (default 500, with 0 disabling and -1 unlimited), and the oldest, lowest-energy workloads are evicted first once the limit is reached; minTerminatedEnergyThreshold defaults to 10 joules. The reason is that a short-lived Job pod must remain visible long enough for Prometheus to scrape its final counter value. With a 30 second scrape interval (the Helm default for the ServiceMonitor) and a Job that completes in 8 seconds, you may never see it unless retention covers the gap. If you run batch or CI workloads, test this explicitly.

Deployment mechanics

The Helm chart is installed from an OCI registry with helm install kepler oci://quay.io/sustainable_computing_io/charts/kepler --namespace kepler --create-namespace. The installation guide states the DaemonSet runs privileged with hostPID: true, with defaults of 100m CPU request and a 400Mi memory limit. That privileged context is worth stating plainly, because the README’s claim of only read-only /proc and /sys access describes what the code needs, not what the default chart grants. If your policy engine rejects privileged pods, plan a tailored security context and verify the readers still function.

The verification command in the README is curl http://localhost:28282/metrics | grep kepler_node_cpu_watts, which also tells you the exporter’s default port. The new liveness and readiness probes in 0.12.0 use a dedicated endpoint, so rollouts and node drains behave more predictably than before.

GPU Power and MIG Attribution in 0.12

GPU energy is the part of the footprint that increasingly dominates AI clusters, and also the part where attribution is least solved by hardware. This section separates what Kepler documents from what it does not, and from what the research literature says about the underlying problem.

What the GPU path measures

GPU monitoring is off by default. The configuration reference places it under experimental.gpu with enabled: false, an idlePower value in watts where 0 means auto-detect, and a dcgmEndpoint URL used for MIG power attribution. Idle power is auto-detected as the minimum observed consumption, which has an obvious consequence: a GPU that has never been idle since the exporter started will have an overestimated idle floor, so active power will be underestimated.

For NVIDIA devices the installation guide says setting config.experimental.gpu.enabled=true deploys an nvidia-libs init container that copies the libnvidia-ml.so* files and sets LD_LIBRARY_PATH to /usr/local/nvidia/lib64, using a 200Mi emptyDir. It also warns that enabling the binary flag without the init path, or the reverse, is not a supported split. Host driver directories such as /run/nvidia/driver are mounted for GPU scenarios. The resulting node series are kepler_node_gpu_watts, kepler_node_gpu_active_watts and kepler_node_gpu_idle_watts plus joule counters, labelled by gpu, gpu_uuid, gpu_name and vendor. Workload series follow: kepler_container_gpu_watts, kepler_pod_gpu_watts and kepler_process_gpu_watts, with matching _joules_total counters.

Why MIG is the hard case

Multi-Instance GPU partitions supported NVIDIA GPUs into as many as seven GPU instances, per NVIDIA’s DCGM documentation, each with dedicated compute and memory slices. DCGM can report metrics at device granularity or at MIG instance granularity. Power is the exception that matters here. A guide from Netdata puts it directly: all instances draw from one enforced power limit, and one instance can be throttled because a neighbour saturated the shared budget. The board has one power sensor.

An IBM Research paper, “On the Partitioning of GPU Power among Multi-Instances” (Vamja, Ray, George and Devi, arXiv 2501.17752), states the problem as the lack of hardware support for per-partition power, with only aggregate GPU power available. Its findings are the best public statement of why naive splitting fails. Power is not additive across concurrent workloads, concurrent FP32 and FP64 operations consumed less power than linear models predicted, and different workloads such as matrix multiplication, LLM inference and GPU Burn have very different power profiles, so a single offline model does not transfer. The authors’ answer is runtime models built from DCGM metrics with MIG-specific features, separate idle and active handling, and scaling to the measured total so aggregate error vanishes.

What Kepler 0.12 adds, and what is not documented

The release note line is “add MIG power attribution via dcgm-exporter”, and the configuration key is dcgmEndpoint. The mechanism is that Kepler takes whole-GPU power from NVML as before, and consults dcgm-exporter for per-instance activity so it can divide the GPU’s active power among the instances and hence among the pods bound to them. Figure 3 sketches that flow.

Kepler MIG power attribution combining NVML board power with dcgm-exporter per-instance activity to produce pod GPU watts

Figure 3: MIG power attribution as inferred from the release notes and configuration keys. The exact weighting function is not documented in the sources I could read.

I have to be explicit about the limits of what I verified. The release notes and the config reference establish that the feature exists, that it depends on a dcgm-exporter endpoint, and that it surfaces as container and pod GPU metrics. I could not find documentation of the exact weighting, for example whether it uses graphics-engine activity, SM activity, memory activity or instance size. Do not assume it matches the IBM paper’s runtime-learning approach. Anything beyond “proportional to some per-instance activity signal supplied by DCGM” would be speculation, and the sketch in Figure 3 is a reading of the interfaces, not of the code.

What the physics implies regardless of the weighting is useful for interpreting the output. Total GPU energy is measured and therefore reliable. The split between two pods sharing a card is a model output with the limits the IBM paper describes. A pod on a 1g.10gb slice running dense FP16 kernels may truly draw more per unit of allocated slice than a neighbour running memory-bound work, and an activity-proportional split will mis-assign that. Use MIG per-pod numbers for ranking and trend, and the whole-card number for budgets.

The time-sliced and whole-GPU cases are simpler. If a pod owns a whole device, its GPU watts equal the device’s active watts, and the attribution is as accurate as NVML itself. Many production inference fleets run this way, so the MIG caveat applies to the subset using partitioning.

Deeper Analysis: Metrics, PromQL and the Carbon Step

Metric families to know

The metrics reference lists node, container, process, VM, pod, build and platform families. For Kubernetes work you mostly use the pod and node families. All joule metrics are counters, so you query them with rate() or increase(); the watt metrics are gauges computed by Kepler itself.

Family Key series Notable labels
Node CPU kepler_node_cpu_watts, _active_watts, _idle_watts, _joules_total, kepler_node_cpu_usage_ratio zone, path
Node GPU kepler_node_gpu_watts, _active_watts, _idle_watts gpu, gpu_uuid, gpu_name, vendor
Pod CPU kepler_pod_cpu_watts, kepler_pod_cpu_joules_total pod_id, pod_name, pod_namespace, state, zone
Pod GPU kepler_pod_gpu_watts, kepler_pod_gpu_joules_total pod_id, pod_name, pod_namespace, state
Container kepler_container_cpu_watts, kepler_container_gpu_watts, kepler_container_cpu_seconds_total container_id, container_name, runtime, pod_id
VM kepler_vm_cpu_watts, kepler_vm_cpu_joules_total vm_id, vm_name, hypervisor
Platform kepler_platform_watts source, node_name, bmc_id, chassis_id
Build kepler_build_info arch, version, revision, goversion

Two absences are worth noting. The metrics reference I read contains no pod-level DRAM or other-component series separate from the zone label, and it contains no carbon metric at all. Kepler produces energy; everything downstream is yours.

PromQL you can use

Average pod power over five minutes, summed across zones to avoid dashboard clutter. Remember the double counting caveat and filter to a single zone if your hardware exposes nested zones:

sum by (pod_namespace, pod_name) (
  rate(kepler_pod_cpu_joules_total{zone="package"}[5m])
)

Because rate() of a joule counter is watts, this equals the pod’s mean CPU package power. Pod energy in kilowatt-hours over a day, by namespace, which is the usual chargeback unit:

sum by (pod_namespace) (
  increase(kepler_pod_cpu_joules_total{zone="package"}[1d])
) / 3.6e6

One kWh is 3.6 million joules, hence the divisor. GPU energy needs a separate query because GPU series have no zone label:

sum by (pod_namespace) (
  increase(kepler_pod_gpu_joules_total[1d])
) / 3.6e6

A conformance check for the attribution is the gap between node and pod power. It should approximate idle power, and it should be stable:

sum(kepler_node_cpu_watts{zone="package"})
-
sum(kepler_pod_cpu_watts{zone="package"})

Compare the result with sum(kepler_node_cpu_idle_watts{zone="package"}). A large unexplained difference points to processes outside Kubernetes pods (system daemons, the kubelet, Kepler itself) which Kepler attributes at process level but not to a pod. Those are real consumers of the active budget that no tenant is charged for.

Turning joules into carbon

Carbon is energy times the carbon intensity of the electricity that supplied it. For location-based accounting the intensity is a grid-average figure in grams of CO2 equivalent per kWh, which varies by region and hour. Kepler does not ship this, so you need a second time series. The usual pattern is an exporter that publishes the intensity for your region as a gauge, either from a grid-data provider or from a static annual factor you maintain yourself.

Kepler carbon metrics pipeline multiplying pod joules by facility overhead and grid carbon intensity

Figure 4: From pod joules to carbon. Kepler supplies only the first box; the overhead factor and intensity are external inputs with their own uncertainty.

With a gauge named grid_carbon_intensity_g_per_kwh (an illustrative name, not a standard one), the pod carbon rate in grams per hour is:

sum by (pod_namespace, pod_name) (
  rate(kepler_pod_cpu_joules_total{zone="package"}[5m])
) / 1000
* on() group_left()
  scalar(grid_carbon_intensity_g_per_kwh{region="eu-north-1"})

Here rate(...) is watts, dividing by 1000 gives kW, and kW times g per kWh gives g per hour. Two further factors are missing from that formula and both matter. First, facility overhead: Kepler measures server-side energy, and the data-centre or cloud provider’s power usage effectiveness (PUE) multiplies it. A PUE of 1.2, used as an illustration, adds 20 percent. Second, the server’s non-CPU consumption: RAPL package and DRAM exclude fans, PSU loss, storage, networking and GPUs unless you add them. The Redfish platform series is the only route inside Kepler to include these.

Market-based accounting, with renewable certificates or power purchase agreements, cannot be derived from a time-series multiplication and does not belong in a dashboard like this. Keep the distinction visible in your report: location-based numbers explain where emissions occur, and market-based numbers are a contractual claim. The earlier post on FinOps and GreenOps carbon-aware scheduling covers how the hourly intensity signal feeds scheduling decisions, which is where this pipeline earns its keep.

A worked example with illustrative numbers

The following numbers are illustrative, not measurements. Suppose a node reports 200 W of package power at 50 percent CPU usage. Kepler splits that into 100 W active and 100 W idle. A pod accounting for 10 percent of node CPU time receives 10 W, or 0.24 kWh over a day. At an assumed grid intensity of 400 g per kWh and a PUE of 1.2, that pod is charged 0.24 x 1.2 x 400, about 115 g of CO2e per day. The node’s idle 100 W produces 2.4 kWh per day that no pod carries, and with the same factors that is about 1.15 kg. In this example the unattributed idle emissions are ten times the pod’s. That ratio is the reason the active-only policy needs an explicit decision about who owns idle.

Running without RAPL, and on virtual machines

Most managed Kubernetes worker nodes are virtual machines, and most hypervisors do not expose RAPL to guests. This is the largest practical limit on Kepler Kubernetes energy accuracy in public cloud. The current docs say Kepler looks for rapl then hwmon; neither exists in a typical cloud VM, so the exporter has no real meter and the options are the fake meter, which is synthetic, or no CPU series at all. Kepler’s VM detection in 0.12.0 recognizes hypervisor processes (VirtualBox, VMware, Xen) so that bare-metal hosts running guests can attribute to VMs; that is the opposite situation, host side rather than guest side.

The legacy line addressed guests with trained models, per the 2024 CNCF post, which describes estimating a VM’s idle power by scaling a per-core figure, for example 160 W across 8 virtual cores giving about 20 W. I did not find that model-based path documented in the 0.10+ docs, and the README says the rewrite is built on sensor readings. If you need numbers on managed cloud nodes, the realistic options are bare-metal instances where the provider allows RAPL, a cloud provider’s own emissions tooling at account level, or a model-based estimate that you label as modelled. Do not present synthetic or modelled values as measurements.

Trade-offs, Gotchas, and What Goes Wrong

The most citable evidence on accuracy is negative. A paper accepted for ICT4S 2025, “Container-level Energy Observability in Kubernetes Clusters” by Pijnacker, Setz and Andrikopoulos, evaluated Kepler experimentally and concluded that the reported energy metrics were not at a satisfactory level, then proposed an alternative tool, KubeWatt, for specific use cases. That evaluation predates the 0.10 rewrite as far as I can tell from its April 2025 date and the project’s release sequence, so it should not be read as a verdict on 0.12. It does show what a careful evaluation looks like, and you should run your own.

Proportional models misattribute by construction. CPU-time share ignores what the CPU did. A pod running AVX-512 loops and a pod spinning on a lock can show the same CPU time and very different power. Kepler’s split is also blind to memory traffic, which DRAM zones capture at node level only. Expect good ranking of heavy versus light pods and poor precision between similar ones.

Idle is excluded, so totals do not reconcile. Pod sums are lower than node power by design. Teams that compare a summed dashboard with a metered rack will think Kepler is wrong. Decide whether idle is distributed by requests, by usage or left as shared overhead, and apply it outside Kepler.

Sampling interval hides short bursts. With a 5 second default monitor interval, processes shorter than that are ignored, and pods that churn quickly show less energy than they consumed. Lowering the interval raises Kepler’s own cost, and Kepler itself is a consumer on every node.

RAPL is a model too. Intel’s counters are derived, not measured at the rails, on many generations, and they cover package and DRAM, not the whole server. AMD and ARM coverage depends on the kernel and platform. The hwmon path is explicitly under the experimental block.

Permissions are a security trade. After the Platypus power side-channel research, CVE-2020-8694 led to kernel changes restricting energy_uj to root. Ubuntu’s advisory rates it CVSS 3.1 5.6 and describes tightening permissions with chmod 400 on those files. Kepler needs to read them, which is part of why the chart runs privileged. Granting a monitoring agent that access on shared multi-tenant nodes deserves a deliberate review.

GPU defaults. The GPU path is experimental and off by default, idle power is auto-detected from the minimum observed value, and MIG attribution depends on an additional dcgm-exporter deployment whose weighting is not publicly specified in the sources I read. A mis-set dcgmEndpoint fails quietly into whole-device numbers, so verify by comparing kepler_pod_gpu_watts for pods sharing a card against the node total.

Version drift. The README says legacy 0.9.x is frozen. Dashboards, alert rules and community recipes written for old metric names (for instance the eBPF-era kepler_container_joules_total) will return no data on 0.10+ and may do so silently. Pin the chart version, and diff /metrics output after every upgrade.

Cardinality. Process metrics carry pid and exe labels; pod metrics carry pod_id and a zone per RAPL zone. On 100 nodes with 100 pods each, that is hundreds of thousands of series before any process-level data. Drop what you do not query.

Practical Recommendations

Treat Kepler as a relative-efficiency instrument first and an accounting instrument second. It is excellent for answering which workloads are expensive, whether a code change moved the needle on a fixed benchmark, and how a rollout changed energy per request. It is weak as a source of audited Scope 2 figures, because the split, the overhead factors and the carbon intensity are each estimates.

Start on a small set of bare-metal nodes with RAPL, validate against a wall or PDU meter at several load levels, and record the error. Only then extend to the fleet. If your nodes are cloud VMs without RAPL, say so in the report, and do not fill the gap with the fake meter.

Define a policy for idle energy before you publish a number. A common, defensible choice is to allocate node idle energy to pods in proportion to their resource requests, which rewards right-sizing, and to show the allocated and unallocated components separately. Use the same pipeline to feed carbon-aware scheduling decisions only after you trust the relative ordering.

For GPU fleets, prefer whole-device allocation where accounting accuracy matters, and use MIG when utilization economics win, accepting that per-pod GPU watts are modelled. Confidential-computing inference adds its own overhead and visibility constraints, which is why the confidential AI inference with TEE and GPU confidential computing discussion is relevant if you meter protected workloads: sensors outside the enclave still see the energy, but per-process telemetry may not.

Checklist:

  • Confirm which meter Kepler selected on each node (rapl, hwmon or none) and alert if a node falls back.
  • Validate against an external meter and publish the observed error alongside the dashboard.
  • Reconcile node versus pod power weekly and attribute the gap explicitly.
  • Sum only one RAPL zone per metric, and never double count package and core.
  • Add PUE and grid intensity as separate, versioned inputs, and label location-based versus market-based.
  • Pin the Kepler chart version and re-check metric names after upgrades.
  • Test that short-lived Jobs appear, given your scrape interval and maxTerminated.
  • Treat experimental features (GPU, hwmon, Redfish) as pilots until validated.

Frequently Asked Questions

What is Kepler and what does it measure?

Kepler is a CNCF sandbox Prometheus exporter that estimates energy consumption for processes, containers, pods, VMs and nodes in Kubernetes. It reads hardware energy counters such as Intel RAPL, reads per-process CPU time from procfs, and attributes active energy to workloads in proportion to CPU time. Output is joule counters and watt gauges, for example kepler_pod_cpu_joules_total. It measures energy, not carbon, which you derive separately with grid intensity data.

Does Kepler still use eBPF?

The versions before 0.10 did, and many articles still describe that design. The project README says the 0.10.0 and later line is a complete rewrite in Go that needs only read-only access to /proc and /sys, and it removes the CAP_SYSADMIN and CAP_BPF capability requirements. Legacy 0.9.x is stated to be frozen. If a tutorial mentions BPF probes or trained model servers, it describes the old line.

Does Kepler work on cloud VMs without RAPL?

Not in the way the legacy versions did. The current configuration walks an ordered list of meters, rapl then hwmon, and uses the first that reports zones. A typical cloud VM exposes neither. The fake meter exists, but it produces synthetic readings for development only. For managed instances, use bare-metal nodes, a provider’s own account-level emissions tools, or clearly labelled modelled estimates instead.

How does Kepler 0.12 attribute power to MIG instances?

The release notes list MIG power attribution via dcgm-exporter, configured with the dcgmEndpoint key under experimental.gpu. Whole-GPU power is measured, and per-instance activity comes from DCGM. The sources I could read do not document the exact weighting function, so treat per-pod MIG watts as modelled. Research from IBM shows power is not additive across concurrent workloads, so expect estimation error.

How do I get carbon from Kepler metrics?

Kepler does not export carbon. Take pod energy in kWh, for example increase(kepler_pod_cpu_joules_total[1d]) / 3.6e6, multiply by a facility PUE, then by grid carbon intensity in grams of CO2e per kWh from a separate time series. Remember CPU package energy excludes fans, power supplies and many other components, so the result is a lower bound on the server’s footprint.

How accurate is Kepler?

There is no single number. Total node energy comes from hardware and is reliable where RAPL exists, while the split among pods is a CPU-time-proportional model. An ICT4S 2025 paper judged an earlier Kepler’s metrics unsatisfactory in its experiments, and that predates the rewrite. Validate against a wall meter on your hardware and workloads, and report the error, rather than relying on a vendor-neutral figure.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *