Software Carbon Intensity (SCI): Measuring and Reducing the Carbon Footprint of Software

Software Carbon Intensity (SCI): Measuring and Reducing the Carbon Footprint of Software

Software Carbon Intensity (SCI): Measuring and Reducing the Carbon Footprint of Software

Most engineering teams can tell you their p99 latency to the millisecond and their cloud bill to the dollar, but cannot tell you whether last quarter’s release made their service cleaner or dirtier. The usual carbon number they do have, a monthly cloud-provider total, rises with traffic and falls with a sunny week on the grid, so it says almost nothing about whether the code improved. Software carbon intensity (SCI) is the metric designed to fix that: a rate, not a total, expressed as grams of CO2-equivalent per unit of useful work.

It matters now because carbon reporting is moving from marketing slides to audit trails, and because AI and batch workloads have made the energy line item large enough to notice. The SCI specification, published by the Green Software Foundation and adopted as ISO/IEC 21031:2024, gives engineers one formula they can compute in a pipeline, compare across releases and defend in a review.

This article explains the formula term by term, shows how to get honest numbers for energy, grid intensity and embodied carbon, walks through a worked example with a runnable calculator, and covers carbon-aware scheduling and the ways the metric can mislead.

What this covers: the SCI formula and its boundary rules, energy measurement options, marginal versus average grid carbon intensity, embodied carbon amortisation, time and location shifting, a Python calculator and scheduler, and the limits and greenwashing risks.

Context and Background

Carbon accounting for IT grew out of corporate reporting. The Greenhouse Gas Protocol splits emissions into Scope 1 (direct), Scope 2 (purchased electricity) and Scope 3 (value chain), and a cloud customer’s footprint lands mostly in Scope 3 because the servers belong to someone else. Cloud providers publish dashboards that allocate those emissions to customers, and open tools such as Cloud Carbon Footprint estimate them from billing and usage data. These are useful for inventory, but they answer a different question from the one a developer asks. A developer wants to know whether a code change helped, and an inventory total cannot tell them, because it moves with demand, region mix and grid weather.

The Green Software Foundation, a Linux Foundation project whose founding members include Microsoft, Accenture, GitHub and ThoughtWorks, set out to define a rate-based metric instead. The result is the Software Carbon Intensity specification. The ISO catalogue lists it as ISO/IEC 21031:2024, “Information technology: Software Carbon Intensity (SCI) specification”, published in March 2024. The specification is deliberately short. It defines a formula, a set of rules about what to include, and a reporting expectation, and leaves measurement methods to the implementer.

Two properties make SCI different from earlier practice. First, it is a score, not an inventory. It divides emissions by a functional unit such as an API call, a user or a transaction, so a service that doubles its traffic with the same infrastructure sees its score fall. Second, it is not an offset-friendly number. The specification requires location-based grid factors and states that market-based instruments, meaning renewable energy certificates, power purchase agreements and offsets, cannot reduce the score. A developer can only lower SCI by using less energy, using energy at cleaner times and places, or using less hardware.

SCI sits next to, not in place of, other methods. Kubernetes teams will meet it alongside per-pod energy attribution, which our Kepler energy attribution guide covers in depth, and finance teams will meet it as the carbon half of a joint cost-and-carbon view, discussed in FinOps and GreenOps for cloud cost and carbon-aware scheduling. Where the numbers get large, as with generative AI inference, the same arithmetic applies, and our trace of AI query energy and water footprint sources shows how fragile the inputs can be.

A final piece of context: the formula has been stable, but tooling and datasets move quickly. Treat provider-specific constants and API behaviour quoted in this article as a snapshot, and check the linked sources before you encode them in a compliance pipeline.

The SCI Formula: What Each Term Means

Direct answer: The SCI score is ((E x I) + M) per R. E is the energy the software consumes in kilowatt-hours, I is the location-based grid carbon intensity in grams of CO2e per kilowatt-hour, M is the share of hardware embodied emissions attributed to the software, and R is the functional unit, such as one request. The result is grams of CO2e per R.

Software carbon intensity formula combining energy, grid intensity, embodied carbon and functional unit into an SCI score

Figure 1: How the SCI score is assembled. Operational emissions (E times I) and embodied emissions (M) are summed, then divided by the functional unit R.

Figure 1 shows the assembly order. Energy and grid intensity multiply to give operational emissions, which the specification writes as O. Embodied emissions M are added, and the sum is divided by R. Because R sits outside the bracket, the score is an intensity: it is the quantity you would plot on a dashboard next to latency per request, and it is what you compare between releases.

Energy (E)

E is the energy consumed by the software system for the functional unit, in kWh. The specification is explicit that it should cover the hardware reserved or provisioned, not merely the hardware that happened to be busy. That choice is a deliberate incentive. If you reserve sixteen cores and use two, the idle fourteen still count, so rightsizing and consolidation show up as SCI improvements. For data centres, E should also reflect facility overhead, usually via power usage effectiveness (PUE), the ratio of total facility energy to IT equipment energy.

In practice E is the hardest term to obtain, and most of the later sections are about how to estimate it without pretending to precision you do not have. The key discipline is to record how E was obtained, whether measured at the socket, read from CPU counters, or modelled from utilisation, because the confidence interval differs by an order of magnitude between those methods.

Region-specific carbon intensity (I)

I is the carbon emitted per kWh of electricity, in gCO2e/kWh, for the grid the software runs on. The specification requires a location-based factor. For grid-connected electricity, it permits short-run marginal, long-run marginal or average intensity and does not prescribe which, so the choice is yours and must be stated. We return to that choice below because it changes both the number and the optimisation incentives.

Embodied emissions (M)

M is the portion of the hardware’s manufacturing, transport and end-of-life emissions attributed to the software. The specification expands it as M = TE x TS x RS, where TE is the total embodied emissions of the hardware (from a life-cycle assessment), TS is the time-share, the fraction of the hardware’s expected lifespan reserved for the software (time reserved divided by expected lifespan), and RS is the resource-share, the fraction of the hardware’s resources reserved for it. A service that holds 4 of a host’s 64 vCPUs for one day of a four-year lifespan is charged 1/1461 of the time and 1/16 of the resources.

Functional unit (R)

R is the unit that describes how the application scales: an API call, a user, a transaction, a build, a minute of video. Every term in the formula must be computed for the same R. The choice is the most consequential design decision in an SCI programme, because it defines what “more useful work” means. A functional unit that tracks real value, such as a completed checkout, rewards genuine efficiency. One that tracks raw activity, such as an HTTP request, can be gamed by splitting one operation into many.

The boundary

The specification asks you to include all supporting infrastructure that contributes significantly to operation, and it names a long list: compute, storage, networking, memory, monitoring, idle machines, logging, scanning, build and deploy pipelines, testing, ML training, backup, redundancy, failover and end-user, IoT and edge devices. The principle is a systems-level view. If you shrink the production service by pushing work into an uncounted batch cluster, the score improves and the planet does not. We treat boundary choice as a first-class decision in the next section.

Why the shape of the formula drives behaviour

The formula has two multiplicative levers (E and I), one additive lever (M) and one divisor (R), and each maps to a distinct engineering action. Reduce E with efficient code, caching, right-sizing and better hardware. Reduce I by moving work in time or place, which is carbon-aware computing. Reduce M by running hardware longer, packing it fuller and buying fewer machines. Raise R by getting more useful work from the same footprint. That mapping is the whole practical value of the metric: it turns “be greener” into four named levers, each with an owner.

Getting Honest Inputs: Boundary, Energy, Grid Intensity and Embodied Carbon

Writing the formula is easy. The work is in feeding it defensible inputs, and in being clear about where each input is measured and where it is merely modelled. This section takes the four inputs in the order a team usually tackles them: boundary first, because it determines what to measure.

Choosing and documenting the boundary

Start by drawing the software boundary on a diagram and listing every component inside it. A typical web service includes application containers, a database, a cache, a message broker, a load balancer, observability agents and the logging pipeline. It also includes the CI system that builds and tests it, the staging environment, backups, and the replicas kept for failover. The specification’s emphasis on idle machines and redundancy matters here: a hot standby that handles zero requests still burns energy and carries embodied carbon, and it belongs in the score.

Two boundary mistakes recur. The first is truncation, where only the production request path is counted because it is the easiest to meter. That flatters the score and hides build, test and ML training costs, which for some products exceed serving costs. The second is double counting across teams, where a shared platform such as a database cluster is charged in full to each tenant. Allocate shared components by a stated key, for example request share or reserved resource share, and write the key down so the number can be reproduced.

A practical technique is to publish the boundary as a table with three columns: component, how E is obtained, and how it is allocated to R. Auditors and teammates can then see which rows are measured and which are estimated, and you can upgrade the estimated rows one at a time.

Measuring energy

There are three families of methods, and Figure 2 shows how they converge on the same energy figure.

Energy measurement paths for software carbon intensity from RAPL counters, utilisation models and power meters into energy per functional unit

Figure 2: Three ways to obtain E. Hardware counters, utilisation-based models and physical meters feed an attribution step, then PUE is applied before dividing by the functional unit.

Hardware counters. On Intel and AMD CPUs, Running Average Power Limit (RAPL) exposes cumulative energy counters. On Linux the powercap framework surfaces them under /sys/class/powercap (the kernel documentation describes the control type at /sys/class/power_cap/intel-rapl). The package zones are named intel-rapl:0, intel-rapl:1 and so on, each with subzones for core and uncore, and each zone has an energy_uj file reporting the counter in microjoules and a max_energy_range_uj file giving the wrap point. Intel RAPL does not provide an instantaneous power reading, so you sample the counter twice and divide the difference by elapsed time. The counters cover the CPU package and, on some platforms, DRAM, but not the whole server: fans, disks, network cards and accelerators are outside them unless measured separately.

Utilisation models. When you cannot read counters, which is the situation inside most public-cloud virtual machines, you estimate. The Cloud Carbon Footprint methodology, derived from Etsy’s Cloud Jewels, computes average watts as minimum watts plus utilisation times the difference between maximum and minimum watts, then multiplies by vCPU hours. Its documentation lists fallback averages per provider when the processor is unknown (for example 0.74 W minimum and 3.5 W maximum per vCPU for AWS), uses the provider-reported utilisation where available and a 50 percent assumption otherwise, and applies per-provider PUE values. Those constants are the tool’s published defaults, not measurements of your workload, and they will be wrong for any specific instance type by an unknown margin.

Physical meters. Smart PDUs, rack power strips and baseboard management controllers (BMCs) report wall power for the whole server. They are the best reference available outside a lab. On a fleet you own, a periodic comparison between PDU readings and RAPL-derived estimates lets you calibrate a scaling factor for the unmetered share of the machine.

Attribution. Whole-machine energy then has to be split among tenants and processes. Kubernetes projects such as Kepler read counters and eBPF-derived CPU time per container to apportion node energy to pods; the trade-offs are covered in our Kepler guide linked above. Whatever tool you use, the split is an allocation model, not an observation, and the SCI report should say so.

Accelerators. GPUs report board power through vendor management libraries, and that figure is reasonably direct. The hard part is sharing: attributing the energy of one multi-tenant GPU to several workloads depends on time slicing or partitioning, and any two attribution schemes can disagree materially.

A worked reminder about PUE: it is a multiplier on IT energy, typically between 1.1 and 1.2 for the large providers’ own published figures, so the direct effect on E is modest. Do not let a good PUE distract from the larger levers of utilisation and grid intensity.

Grid carbon intensity: average versus marginal

I is where the specification gives you a choice and where the engineering consequences are subtle.

Average intensity is total grid emissions divided by total generation in a period. It answers “what is the carbon footprint of the electricity mix right now”, it is what most public datasets publish, and it is the right input for accounting the footprint of electricity you consumed. Marginal intensity is the emissions change caused by one additional unit of demand, which usually means the plant that ramps up at the margin, often gas. It answers “what happens to emissions if I run this job”, which is the right input for a decision about whether to shift load.

The two can diverge sharply. At midday on a grid with abundant solar, average intensity may be low while the marginal plant is still a gas turbine, or while curtailed solar makes the true marginal intensity near zero. At night average intensity can be moderate while the marginal plant is coal or gas. A scheduler optimising on average intensity will chase the visibly green hours; a scheduler using marginal data will chase the hours when extra demand actually displaces dirty generation or absorbs surplus clean energy. The specification does not tell you which to prefer, so the credible approach is to pick one, label it in every report, and avoid mixing it with the other in the same trend line.

Granularity matters as much as type. Annual national averages are fine for a rough order-of-magnitude estimate, but they hide the hour-to-hour and region-to-region variation that carbon-aware computing exploits. Sources include real-time and forecast services such as Electricity Maps and WattTime, public datasets like the US EPA eGRID (which the Cloud Carbon Footprint tool uses for AWS and Azure regions in the United States), and cloud-provider region figures; Google publishes per-region carbon data, and CCF can adjust it by carbon-free energy percentage. Each source has different methodology, lags and licensing, so record the source and timestamp alongside every I value.

Embodied carbon

Hardware carries a carbon debt from mining, fabrication, assembly and shipping before it draws a single watt. For servers, vendor product carbon footprint sheets or published life-cycle studies give the total embodied emissions (TE), usually in kgCO2e per server. The SCI formula amortises TE over the expected lifespan (EL) and charges your software for the time it reserves (TiR) and the fraction of the machine it reserves (RR over ToR).

Two things follow. First, lifespan is a lever with real teeth: extending the expected life of a fleet from four to six years reduces the time-share charge per hour by a third, with no code change. Second, reservation, not usage, drives M, which again rewards right-sizing: an oversized reservation pays the embodied charge for resources it never touches. For cloud instances you rarely see TE directly, so teams use provider-published or third-party estimates, and CCF’s methodology page notes that its embodied estimates apply to compute usage only. Be especially careful not to ignore M for end-user and edge devices if they fall inside your boundary: for a mobile app the phone’s embodied carbon can dominate the operational energy of the backend.

A Worked Example and a Runnable SCI Calculator

The numbers below are illustrative, chosen to show the arithmetic and the sensitivity, not measurements of any real service or grid.

Suppose a service reserves 4 of 64 vCPUs on a host for a 24-hour window. The host’s total embodied emissions are taken as 1,200 kgCO2e (1,200,000 g) with a four-year expected lifespan. In that window the service consumed 3.2 kWh including PUE overhead and handled 250,000 requests. The grid intensity was 420 gCO2e/kWh.

Operational emissions are 3.2 x 420 = 1,344 g. Embodied emissions are 1,200,000 x (24 / 35,040) x (4 / 64), which is about 51.4 g. The total is about 1,395.4 g over 250,000 requests, so the SCI is about 5.58 mg CO2e per request. In the illustration the embodied share is under 4 percent of the total; on a clean grid it matters far more.

Rerun with the same workload on a grid at 90 gCO2e/kWh: operational drops to 288 g, the total to about 339.4 g, and the SCI to about 1.36 mg per request, a 76 percent fall with no code change, while M has grown to 15 percent of the total. This is the key structural insight: as the grid decarbonises, embodied carbon becomes the dominant term, and the optimisation priority shifts from “use less energy” to “use less hardware for longer”.

The calculator below implements the formula and the M expansion, with input validation. It is plain standard-library Python.

from dataclasses import dataclass
from typing import Sequence


@dataclass(frozen=True)
class Embodied:
    te_g: float    # total embodied emissions of the hardware, gCO2e
    el_h: float    # expected lifespan, hours
    tir_h: float   # time reserved for the software, hours
    rr: float      # resources reserved (e.g. vCPUs)
    tor: float     # total resources available (e.g. host vCPUs)

    def m(self) -> float:
        if self.el_h <= 0 or self.tor <= 0:
            raise ValueError("lifespan and total resources must be positive")
        return self.te_g * (self.tir_h / self.el_h) * (self.rr / self.tor)


def sci(energy_kwh: float, intensity_g_per_kwh: float, m_g: float, r: float) -> float:
    """SCI = ((E * I) + M) per R, in gCO2e per functional unit."""
    if r <= 0:
        raise ValueError("functional unit count must be positive")
    return (energy_kwh * intensity_g_per_kwh + m_g) / r


def sci_windows(windows: Sequence[tuple], m_total_g: float) -> float:
    """windows: (energy_kwh, intensity, requests) per interval.
    Operational emissions are summed per interval, so intensity varies by hour."""
    total_r = sum(w[2] for w in windows)
    o = sum(e * i for e, i, _ in windows)
    return (o + m_total_g) / total_r


if __name__ == "__main__":
    m = Embodied(te_g=1_200_000, el_h=4 * 365 * 24, tir_h=24, rr=4, tor=64).m()
    print(round(m, 2), "g embodied")
    print(round(sci(3.2, 420, m, 250_000) * 1000, 3), "mgCO2e per request")
    print(round(sci(3.2, 90, m, 250_000) * 1000, 3), "mgCO2e per request, clean grid")

Running it prints 51.37 g embodied, 5.581 mg per request, and 1.357 mg per request on the cleaner grid, matching the hand calculation. Note sci_windows: a day-long average intensity applied to the whole day’s energy is only correct if load was flat. Real services peak with demand, and demand often correlates with grid intensity, so multiplying hourly energy by hourly intensity and summing is more accurate than multiplying the totals.

Reading energy from RAPL

The next snippet measures the energy of a function on a Linux host that exposes Intel RAPL. It handles counter wrap-around once, which is enough for short measurements; a long-running sampler should poll more often than the wrap interval.

import time
from pathlib import Path

ZONE = Path("/sys/class/powercap/intel-rapl:0")


def read_uj(zone: Path = ZONE) -> int:
    return int((zone / "energy_uj").read_text())


def measure(fn, zone: Path = ZONE):
    wrap = int((zone / "max_energy_range_uj").read_text())
    e0, t0 = read_uj(zone), time.perf_counter()
    fn()
    e1, t1 = read_uj(zone), time.perf_counter()
    delta = e1 - e0 if e1 >= e0 else e1 + wrap - e0   # tolerate one wrap
    return delta / 3.6e12, t1 - t0                    # microjoules -> kWh, seconds


if __name__ == "__main__":
    kwh, secs = measure(lambda: sum(i * i for i in range(10_000_000)))
    print(f"{kwh:.9f} kWh over {secs:.2f} s")

The conversion divides microjoules by 3.6 x 10^12 because one kWh is 3.6 million joules, or 3.6 x 10^12 microjoules. The measurement includes everything running on that CPU package during the window, not just your function, so run it on a quiet machine, repeat it, and subtract an idle baseline. Reading energy_uj has been restricted to privileged users on many kernels after research showed RAPL could be used as a side channel, so expect to need elevated permissions or a vendor-provided exporter. I have not run this snippet against real hardware here, and many virtual machines do not expose the zone at all.

Carbon-Aware Computing: Shifting Work in Time and Place

Once SCI is measured, the I term is the cheapest lever to pull for flexible workloads. Carbon-aware computing means running work when and where the grid is cleaner. The Green Software Foundation’s Carbon Aware SDK is an MIT-licensed project that wraps this idea in a WebApi and a command-line interface with equivalent functionality. According to its repository, it abstracts third-party data providers (it names WattTime and Electricity Maps) and normalises their differing units to gCO2/kWh. The repository recommends the WebApi for larger organisations needing central management and auditability, and the CLI for legacy integration and non-cloud deployments. I could not confirm from the repository front page which providers are currently supported or the exact endpoint list, so check its docs before building against it.

Carbon-aware scheduling decision flow choosing between time shifting, location shifting or running immediately based on job flexibility and forecast intensity

Figure 3: A carbon-aware scheduling decision. Jobs with slack are shifted in time or place using a forecast; inflexible jobs run immediately, and actual intensity is recorded for reporting.

Time shifting

Time shifting delays flexible work, such as batch analytics, model training, media transcoding, backups and CI pipelines on non-urgent branches, to a window with lower forecast intensity. The mechanism needs three inputs: the job’s runtime, its deadline, and an hourly forecast. The scheduler slides a window of the job’s length across the forecast, discards any window that would finish after the deadline, and picks the one with the lowest mean intensity.

from datetime import datetime, timedelta


def best_window(forecast, runtime_h, deadline):
    """forecast: ascending list of (hour_start: datetime, gco2_per_kwh).
    Returns (start, mean_intensity) of the cleanest runtime_h-hour window
    that finishes by `deadline`, or None if nothing fits."""
    best = None
    for i in range(len(forecast) - runtime_h + 1):
        chunk = forecast[i:i + runtime_h]
        end = chunk[-1][0] + timedelta(hours=1)
        if end > deadline:
            break
        mean = sum(v for _, v in chunk) / runtime_h
        if best is None or mean < best[1]:
            best = (chunk[0][0], mean)
    return best


if __name__ == "__main__":
    t0 = datetime(2026, 10, 10, 18)
    vals = [380, 360, 340, 300, 260, 220, 200, 190, 210, 250, 300, 340]  # illustrative
    fc = [(t0 + timedelta(hours=i), v) for i, v in enumerate(vals)]
    print(best_window(fc, 3, t0 + timedelta(hours=12)))

With the illustrative forecast above, a three-hour job with a twelve-hour deadline is moved to start at midnight with a mean of 200 gCO2e/kWh, against 360 for an immediate start, a 44 percent reduction in the operational term for that job. Real forecasts are noisy, so the benefit you realise depends on forecast skill. Log the forecast value you scheduled against and the actual intensity you got; the gap is your forecast error and tells you how much slack is worth paying for.

Time shifting has hidden costs. It adds latency to the shifted work, concentrates load into clean windows (which can create contention, and cost spikes if you use on-demand capacity), and moves a job’s energy rather than removing it. It also interacts with E: a job that waits holds reserved capacity or queue state. Savings are therefore largest for work that is truly deferrable and runs on capacity you would otherwise leave idle.

Location shifting

Location shifting runs work in the region with the cleanest grid, subject to data residency, latency, egress cost and capacity. It suits training runs, large batch jobs and anything with weak locality. It suits interactive services badly, because the latency penalty is borne by users, and it can run into legal limits on where personal data may be processed.

Two cautions apply. The cleanest region in annual average terms is not always the cleanest hour by hour, so a static “always use region X” policy captures only part of the available benefit. And moving data has a carbon cost of its own: network transfer energy is small per gigabyte but not zero, and replicating a large dataset to chase a lower intensity can wipe out the saving. CCF’s published coefficient for inter-data-centre transfer is 0.001 kWh per GB, which is a default estimate and not a universal constant.

Demand shaping

A third pattern, sometimes called demand shaping, adjusts quality of service to grid state without moving the work. A video platform might lower default resolution during high-intensity hours, or an app might postpone background sync. It reduces E and pairs well with SCI, because both the numerator drops and, if you pick R carefully, the unit of work remains meaningful. Document these degradations; shaping that users notice is a product decision, not just an infrastructure one.

Where carbon-aware computing pays and where it does not

The benefit scales with the spread of intensity across the shift options, and that spread varies enormously between grids. A grid dominated by hydro or nuclear has a flat profile with little to gain from shifting; a grid with heavy solar and gas swings widely across the day. The honest way to size the prize is to take a month of your own hourly data, compute the best-case saving for your deferrable workloads using a function like best_window, and compare it with the engineering and latency cost. Many teams find that rightsizing and idle reduction (the E and M levers) deliver more, with less risk, than shifting.

Putting SCI Into a Pipeline

A one-off calculation is useful for a design review. A score that moves with every release is useful for engineering. The architecture in Figure 4 shows a minimal reporting pipeline.

Software carbon intensity reporting pipeline where requests, energy, grid intensity and hardware inventory feed an SCI calculator and report store

Figure 4: A reporting pipeline. Independent feeds for requests, energy, grid intensity and hardware inventory converge on a calculator that stores both the score and its inputs.

Design points

Store inputs, not just scores. Persist E, I, M and R per window with a source tag for each (measured, modelled, estimated) and the grid intensity type. When a methodology changes you can recompute history, and when someone challenges a number you can show the derivation. A score without its inputs cannot be audited.

Align time windows. Energy, intensity and request counts must be sampled over the same intervals. Misaligned windows silently skew the result, particularly for bursty traffic. Fifteen-minute or hourly windows are a common compromise between data availability and resolution.

Treat R as a contract. Define the functional unit once, in code review, and version it. If a team redefines “request” to exclude health checks mid-quarter, the trend line breaks. Use a stable, business-meaningful unit where possible and report a secondary technical unit alongside it.

Gate on regression, not on absolute values. Because input quality differs across services, comparing absolute SCI between teams is risky. A CI check that compares a candidate release against the previous one on the same benchmark, in the same region with the same intensity assumption, isolates the effect of the code. In a load test you can fix I and M, and track E per R, which makes the check deterministic enough to fail a build.

Report uncertainty. If E is modelled from utilisation with default coefficients, state a range. A score of 5.6 mg with a plausible spread of plus or minus 30 percent is more honest, and more useful, than 5.581 mg.

The learning-to-action loop

The best use of SCI is as a prioritisation tool. Decompose each service’s score into O and M shares and into the sub-terms. If M is 40 percent of a service on a clean grid, hardware efficiency and lifespan dominate. If O is 95 percent on a dirty grid, look at utilisation and shifting. If R is rising slower than E, look at what work is unproductive: retries, over-fetching, chatty polling, or oversized responses. That decomposition turns a single number into a ranked backlog.

Trade-offs, Gotchas, and What Goes Wrong

SCI is a good metric with sharp edges. Knowing them is the difference between a useful signal and a greenwashing liability.

Offsets cannot help, and that is the point, but it surprises people. Teams used to corporate reporting expect renewable certificates and power purchase agreements to cut their number. Under the specification they do not. An SCI score is not a net-zero claim, and it should not be described as one. Mixing SCI scores with market-based claims in the same sentence invites exactly the criticism the specification tries to avoid.

The functional unit can be gamed. Optimising a per-request score by splitting one operation into several cheap requests, or by counting synthetic health checks, improves the intensity without reducing the footprint. Tie R to user-visible value, and watch absolute energy alongside the intensity. An intensity metric can fall while total emissions rise, if the service grows faster than it improves: this is the rebound effect, and SCI alone does not show it.

Intensity improvements do not equal absolute reduction. SCI deliberately normalises for scale. Organisations with absolute emission targets still need an inventory view. Treat SCI as the engineering metric and inventory reporting as the accountability metric; they answer different questions.

Modelled energy is not measured energy. Most cloud SCI scores rest on utilisation models with default coefficients. These are acceptable for trends within one service, and poor for comparing across providers or instance families. If a vendor claims a precise per-request figure, ask how E was obtained.

Average and marginal intensity can point in opposite directions. Optimising on one while reporting on the other produces confusing results. Pick one for decisions and say which.

Embodied data is thin. Life-cycle assessments for specific servers, accelerators and cloud instance types are often unavailable or proprietary, and published estimates vary. Treat M as an order-of-magnitude figure, and do not claim a precision the data cannot support. For rapidly evolving AI hardware this is especially true.

Shifting has second-order effects. If many organisations chase the same cleanest hour or region, they can shift demand enough to change the grid’s marginal plant. At today’s scale this is theoretical for most teams, but it is a reason to prefer marginal signals and to avoid assuming savings scale linearly.

Data licensing and availability. Real-time and forecast intensity services have terms of use and rate limits, and coverage varies by country. A scheduler that fails closed when the feed is down will stall batch work; one that fails open will quietly run with stale data. Choose and test the fallback.

Verification. The specification describes a score; it is not itself an audit regime. A self-reported SCI is only as credible as its documented boundary, sources and method. Publish them.

Comparing Approaches: A Decision Matrix

The table compares the main ways teams quantify or reduce software carbon. The ratings are qualitative judgements, not measurements.

Approach Best for Energy data quality Grid data Embodied carbon Main weakness
SCI score per release Engineering regressions, design reviews Depends on method used Location-based required Included by formula Needs a well-chosen R and boundary
Cloud provider carbon dashboards Inventory and Scope 3 reporting Provider-modelled Often market-based or mixed Provider-specific Total, not rate; hard to see code effects
Cloud Carbon Footprint style estimates Multi-cloud inventory from billing and usage Modelled with default coefficients Public datasets such as eGRID Compute-only estimate Coarse for any one workload
Per-pod counter attribution (Kepler style) Kubernetes clusters you control Counter-based where exposed Needs separate feed Not covered Allocation model; not in most managed VMs
Carbon-aware scheduling Deferrable batch, training, CI Unchanged Forecast or marginal feed Unchanged Benefit depends on grid spread and forecast skill
Rightsizing and utilisation work Almost every fleet Reduces E directly Unchanged Reduces M Needs sustained ownership

Reading the matrix: SCI is the common yardstick, and the other rows are either ways to feed it or ways to move its terms. For most teams the sequence that works is to adopt SCI for one service, improve the E and M inputs, take the easy utilisation wins, and only then add shifting for the subset of workloads that can tolerate it.

Practical Recommendations

Start small and make the number reproducible before you make it good. Pick one service with a clear owner, define R as a business-meaningful unit, and write the boundary down as a table. Compute a first SCI from modelled energy and an average intensity for your region, label every input as measured, modelled or estimated, and publish the result with its uncertainty. A rough but documented number beats a precise-looking one nobody can trace.

Then work the levers in order of leverage and risk. Fix utilisation and idle capacity first, because they reduce both E and M and carry no latency cost. Extend hardware lifespan and consolidate where reliability allows. Only then introduce carbon-aware scheduling for workloads with real slack, beginning with CI and batch, and measure the realised saving against forecast skill. Keep reporting separate from offsets, and keep an absolute-emissions view next to the rate.

  • Define R once, version it, and tie it to user value.
  • Document the boundary: include idle, redundancy, CI and storage.
  • Record the source and type (measured, modelled, estimated) of every E, I and M input.
  • Choose average or marginal intensity per use case and label it.
  • Store inputs with each score so history can be recomputed.
  • Gate releases on SCI regression under fixed load, not on absolute values.
  • Rightsize and extend hardware life before shifting workloads.
  • Never describe an SCI score as a net-zero or offset-adjusted claim.
  • Track absolute energy beside the rate to catch rebound effects.
  • Re-verify provider constants and dataset versions at least annually.

Frequently Asked Questions

What is software carbon intensity?

Software carbon intensity (SCI) is a rate-based metric that expresses the carbon emissions of a software system per unit of useful work. It is calculated as ((E x I) + M) per R: energy times location-based grid intensity, plus the share of hardware embodied carbon, divided by a functional unit such as one API call. It is published by the Green Software Foundation and standardised as ISO/IEC 21031:2024.

How do you calculate the SCI score?

Measure or estimate the energy your software uses in kWh, multiply by the grid carbon intensity in gCO2e per kWh to get operational emissions, then add the embodied emissions attributed to your share of the hardware. Divide the total by your functional unit count. Embodied emissions are the hardware’s total embodied carbon times your time-share and resource-share. The runnable calculator above implements each step.

Can renewable energy certificates or offsets lower my SCI score?

No. The specification requires location-based grid factors and states that market-based instruments such as renewable energy certificates, power purchase agreements and offsets cannot reduce the score. The design intent is that developers can only improve SCI by using less energy, using energy when and where the grid is cleaner, or using less hardware. Offsets can still matter for corporate reporting, but they sit outside this metric.

What is the difference between marginal and average grid carbon intensity?

Average intensity is total grid emissions divided by total generation, describing the mix of electricity in a period. Marginal intensity is the emissions change caused by one extra unit of demand, typically set by the plant that ramps up. The SCI specification allows short-run marginal, long-run marginal or average values without prescribing one. Average suits footprint accounting; marginal is generally better for deciding whether to shift load.

What is carbon-aware computing?

Carbon-aware computing runs flexible workloads when and where grid carbon intensity is lower, by time shifting to a cleaner window or location shifting to a cleaner region. Tools such as the Green Software Foundation’s Carbon Aware SDK expose forecast data through a WebApi and CLI. It lowers the I term of SCI, works best for deferrable batch, training and CI jobs, and depends on forecast quality and the variability of the local grid.

Does SCI include the carbon cost of manufacturing servers?

Yes. The M term captures embodied emissions: the hardware’s total life-cycle emissions multiplied by the fraction of its lifespan and the fraction of its resources reserved for your software. Because it is based on reservation, oversized allocations are penalised. As grids get cleaner the embodied share grows, which is why extending hardware life and avoiding idle provisioned capacity become increasingly important levers.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *