Differential Privacy for Machine Learning: DP-SGD, Privacy Budgets and What Epsilon Really Means

Differential Privacy for Machine Learning: DP-SGD, Privacy Budgets and What Epsilon Really Means

Differential Privacy for Machine Learning: DP-SGD, Privacy Budgets and What Epsilon Really Means

A trained model is a compressed, queryable copy of its training data, and it leaks. Language models can emit memorised training strings, classifiers answer more confidently on rows they have seen, and an attacker with nothing but API access can often tell whether a specific person’s record was in the training set. Removing names from the data does not fix this, because the leak is in the weights, not in the column headers. Differential privacy machine learning is the one defence with a mathematical guarantee: it bounds how much any single record can change what the training algorithm outputs.

The guarantee comes with a dial, epsilon, and the dial is where most teams go wrong. A budget of epsilon equal to 1 and a budget of epsilon equal to 50 both get called “differentially private” in slide decks, yet they promise very different things. Choosing and reporting epsilon honestly is the hard part, not running the library.

This article explains the definition, the Laplace and Gaussian mechanisms, DP-SGD, the subsampled Gaussian and the accountants that track it, then runs a real Opacus experiment and ends with a plain-language reading of epsilon.

What this covers: the epsilon-delta definition, noise calibration, the DP-SGD algorithm step by step, RDP and PRV accounting, a runnable Opacus sketch with measured numbers, membership-inference attacks, DP fine-tuning of language models, utility costs, how to choose epsilon, and a decision matrix against federated learning, anonymization and synthetic data.

Context and Background

Differential privacy (DP) was introduced by Cynthia Dwork, Frank McSherry, Kobbi Nissim and Adam Smith in 2006. Its premise is a change of question. Instead of asking “is this dataset anonymous?”, which depends on what an attacker already knows, it asks “how much does the output of this computation depend on any one person’s data?”. An algorithm is private if removing or changing one record barely moves the distribution of its outputs. The standard reference text is Dwork and Roth, “The Algorithmic Foundations of Differential Privacy” (2014), and the machine learning story starts from Abadi et al., “Deep Learning with Differential Privacy” (2016), which introduced DP-SGD and the moments accountant. The paper’s abstract claims that deep networks with non-convex objectives can be trained “under a modest privacy budget”, at a manageable cost in software complexity, training efficiency and model quality (the full text is on arXiv 1607.00133).

The reason ML needs a formal guarantee is the history of failed informal ones. Shokri, Stronati, Song and Shmatikov showed in “Membership Inference Attacks against Machine Learning Models” (2017) that a black-box attacker can train a model to distinguish training from non-training records, including on models from commercial machine-learning-as-a-service providers. Later work, notably the likelihood-ratio attack by Carlini et al. (“Membership Inference Attacks From First Principles”, 2022), showed that evaluating attacks by average accuracy understates the risk: what matters is the true positive rate at a very low false positive rate. Those results frame everything below. A defence is only as good as the attacks it has been measured against, and DP is the one defence whose guarantee holds against every attack, including ones nobody has invented yet.

Government statistical agencies were early adopters at scale. The US Census Bureau applied differential privacy to 2020 Census products such as the Redistricting Data Summary File, adding calibrated noise to published counts. Its own disclosure avoidance page now notes that the Bureau is evaluating alternatives following a Department of Commerce order prohibiting noise infusion, which is a useful reminder that the accuracy cost of DP is a policy fight, not just a technical detail. Industry has deployed local differential privacy for telemetry collection. I have not verified the specific parameters of those deployments for this post, so I do not quote them.

For practitioners who already know federated learning, one correction is worth stating early. Federated learning keeps raw data on devices, but the shared model updates still leak. FL is a data-movement strategy; DP is a leakage bound. They compose, and the section on trade-offs explains how. The practical DP toolchain today is Opacus for PyTorch, plus JAX and TensorFlow Privacy implementations of the same algorithm, and it is mature enough to apply to fine-tuning real models.

The Core Argument: What the Guarantee Says and How DP-SGD Delivers It

Direct answer: A randomized algorithm M is (epsilon, delta)-differentially private if, for any two datasets differing in one record and any set of outputs S, the probability that M outputs something in S changes by at most a factor of e to the epsilon, plus a small slack delta. DP-SGD achieves this for training by clipping each example’s gradient and adding Gaussian noise before every update.

Differential privacy machine learning definition showing two neighbouring datasets, a randomized mechanism and the bounded output ratio

Figure 1: The epsilon-delta guarantee. Two neighbouring datasets produce output distributions whose ratio is bounded by e to the epsilon, apart from a delta of failure probability.

Figure 1 is the whole definition in one picture. Take a dataset D and a neighbour D’ that differs by one record, either added or removed or substituted depending on the convention. Run the mechanism on both. The two output distributions must be nearly indistinguishable, which formally means Pr[M(D) in S] is at most e^epsilon times Pr[M(D’) in S] plus delta, for every measurable set S. If epsilon is small, observing the output tells an adversary almost nothing about whether any particular person participated. Delta is the probability that the pure epsilon guarantee fails outright; it should be smaller than the reciprocal of the dataset size, with 1e-5 a common choice for datasets of tens of thousands of rows. A delta around 1/n would allow a mechanism that publishes one random record, which is plainly not private.

The building blocks: sensitivity, Laplace and Gaussian noise

Every DP mechanism rests on one quantity, the sensitivity of the function being computed. The L1 sensitivity of a query f is the largest change in f’s output when one record changes, and the L2 sensitivity is the same in the Euclidean norm. A counting query has sensitivity 1. A sum of values each clipped to [0, B] has sensitivity B. This is why clipping sits at the heart of DP-SGD: it forces a bounded sensitivity on a quantity, the gradient, that is otherwise unbounded.

The Laplace mechanism adds noise drawn from a Laplace distribution with scale equal to the L1 sensitivity divided by epsilon, and gives pure epsilon-DP with delta equal to zero. The Gaussian mechanism adds noise with a standard deviation scaled to the L2 sensitivity, and gives (epsilon, delta)-DP. The classical calibration, from Dwork and Roth Theorem A.1, requires a standard deviation of at least sensitivity times the square root of 2 ln(1.25/delta), divided by epsilon, and is valid for epsilon below 1. For delta equal to 1e-5 the multiplier square root of 2 ln(1.25/delta) is about 4.84, so a single Gaussian release at epsilon 1 needs noise nearly five times the sensitivity. That classical bound is loose and limited to small epsilon, and the analytic Gaussian mechanism and the accountants below are tighter. Deep learning uses the Gaussian rather than the Laplace for a practical reason: Gaussian noise composes cleanly under Renyi accounting and its L2 geometry matches gradient clipping by L2 norm.

DP-SGD step by step

Standard stochastic gradient descent computes an average gradient over a mini-batch and steps along it. That average is a function of the whole batch, and one outlier example can dominate it. DP-SGD changes three things, shown in Figure 2.

DP-SGD differential privacy machine learning loop with Poisson sampling, per-sample clipping, Gaussian noise and accountant logging

Figure 2: One DP-SGD iteration. Sample, compute per-example gradients, clip each, sum, add Gaussian noise scaled to the clip norm, normalize, step, and tell the accountant what happened.

First, gradients are computed per example, not per batch. Second, each per-example gradient is clipped so that its L2 norm is at most C, the clipping norm: the gradient g becomes g times min(1, C / ||g||). After clipping, any one example can move the batch sum by at most C, so the sum has L2 sensitivity C. Third, Gaussian noise with standard deviation sigma times C is added to the summed gradient, where sigma is the noise multiplier. The noisy sum is divided by the expected batch size and passed to an ordinary optimizer. Because the privacy argument concerns only the noisy gradient, the optimizer can be SGD, momentum or Adam without any change to the guarantee; this is the post-processing property, which says that anything computed from a private output stays private.

The sampling step matters as much as the noise. Batches are formed by Poisson sampling: each row independently joins a batch with probability q, the sampling rate. This randomness is a source of privacy in itself. A record that is absent from most batches only influences a small fraction of steps, and the privacy analysis exploits that amplification by subsampling. The consequence for implementers is that the standard practice of shuffling the data and slicing fixed-size batches does not match the analysis, and an epsilon computed under the Poisson assumption does not formally apply to it. Opacus uses Poisson sampling by default for this reason (its poisson_sampling argument defaults to true).

Why composition makes the budget leak away

A single noisy release is easy to analyse. Training is thousands of releases, and each reveals a little. Composition theorems say how the losses add up. Basic composition adds epsilons linearly, so T steps at epsilon each cost T times epsilon, which is useless for deep learning. Advanced composition gives roughly the square root of T growth, and the Renyi analysis described below does better still. The concrete numbers in the next section show that without subsampling, even tight accounting makes deep learning impractical, and with subsampling it becomes possible.

The subsampled Gaussian and Renyi differential privacy

The object DP-SGD actually analyses is the subsampled Gaussian mechanism: sample with rate q, then apply a Gaussian mechanism with noise multiplier sigma. Abadi et al. introduced the moments accountant to track its privacy loss through composition. Mironov generalised that idea as Renyi differential privacy (RDP) in 2017 (arXiv 1702.07476), a definition based on Renyi divergence of order alpha that composes by simple addition. A mechanism is (alpha, epsilon_alpha)-RDP if the Renyi divergence of order alpha between its outputs on neighbouring datasets is at most epsilon_alpha.

For the plain Gaussian mechanism with sensitivity 1 and noise multiplier sigma, the RDP curve is epsilon_alpha equal to alpha divided by 2 sigma squared. After T identical, non-subsampled steps the curves simply add, giving T alpha / (2 sigma squared). RDP is then converted to the (epsilon, delta) form with epsilon equal to epsilon_alpha plus ln(1/delta) divided by (alpha – 1), minimised over alpha; tighter conversions exist and newer libraries use them. The subsampled Gaussian has a more complicated, but computable, RDP curve (Mironov, Talwar and Zhang, 2019, “Renyi Differential Privacy of the Sampled Gaussian Mechanism”). The next section puts numbers on all of this.

Privacy Accounting in Practice: Numbers, Accountants and an Opacus Run

Direct answer: Privacy accounting is the bookkeeping that converts a training run’s noise multiplier, sampling rate and step count into a final (epsilon, delta). Opacus offers RDP, Gaussian-DP and PRV accountants, with PRV the default in current releases. You either fix the noise and read off epsilon, or fix a target epsilon and let the library solve for the noise.

Privacy accounting for DP-SGD comparing RDP and PRV accountants feeding an epsilon delta budget check

Figure 3: The accounting loop. Every step reports its noise multiplier and sampling rate; an accountant composes them and converts the result to (epsilon, delta) so training can stop when the budget is spent.

Why subsampling is the whole game

Consider the plain Gaussian mechanism with noise multiplier sigma equal to 1 and no subsampling, using the RDP formula above and delta equal to 1e-5. I minimised T alpha / 2 + ln(1/delta) / (alpha – 1) over alpha numerically. One release gives epsilon of about 5.3. One hundred releases give about 98. One thousand releases give about 652. Those are the figures for training with full-batch gradient steps, and they are far beyond any meaningful guarantee.

Now compare a subsampled run. With a dataset of 10,000 rows, a batch of 256 (so q is 0.0256), five epochs (about 195 steps) and sigma equal to 1.0, the Opacus RDP accountant reports epsilon of about 2.76 at delta 1e-5. That is roughly 195 noisy steps for a cost that full-batch noise would exceed after fewer than ten. The table below shows what the accountant reports as sigma varies for that same setup; I generated it by running RDPAccountant from Opacus 1.6.0 in a sandbox, and the values are specific to this dataset size and schedule.

Noise multiplier sigma Epsilon at delta 1e-5 (q = 0.0256, about 195 steps)
0.6 10.2
0.8 4.76
1.0 2.76
1.5 1.25
2.0 0.82

A larger-scale schedule shows the other pressure. With 60,000 records, a batch of 1,024 (q about 0.017) and 30 epochs, sigma of 0.8 gives epsilon near 7.96, sigma of 1.1 gives about 4.02, and sigma of 1.5 gives about 2.44. Epsilon grows with the number of steps and with q, and shrinks with sigma. Every extra epoch is paid for in budget, which means DP training pushes you toward fewer epochs, larger batches and better initialisation, exactly the recipe that DP fine-tuning later adopts.

RDP versus PRV accountants

The RDP accountant is easy to implement and composes exactly in the RDP domain, but the final conversion to (epsilon, delta) loses some tightness. The PRV (privacy random variable) accountant, introduced by Gopi, Lee and Wutschitz in “Numerical Composition of Differential Privacy” (2021), works with the full distribution of the privacy loss. It computes the composition numerically using fast Fourier transforms and yields a tighter epsilon with an error estimate. Opacus documents accountant choices of rdp, gdp and prv, with prv as the default in PrivacyEngine. Which accountant you use changes the reported epsilon for identical training, so a published epsilon without the accountant name and sampling method is incomplete reporting. Treat the accountant like a unit of measure.

A runnable Opacus experiment

The script below trains a small classifier on a synthetic five-class Gaussian-blob problem with deliberately overlapping classes, once without privacy and then with target epsilons of 0.5, 1, 3 and 8. It uses make_private_with_epsilon, which solves for the noise multiplier given a target epsilon, delta and epoch count. I ran it with Opacus 1.6.0 and PyTorch 2.14 on CPU.

import torch, torch.nn as nn
from torch.utils.data import DataLoader, TensorDataset
from opacus import PrivacyEngine

torch.manual_seed(0)

def make_data(n, d=20, k=5):
    g = torch.Generator().manual_seed(1)
    centers = torch.randn(k, d, generator=g) * 0.35
    y = torch.randint(0, k, (n,), generator=g)
    x = centers[y] + torch.randn(n, d, generator=g)
    return x, y

xtr, ytr = make_data(10000)
xte, yte = make_data(2000)
loader = DataLoader(TensorDataset(xtr, ytr), batch_size=256)

def run(target_eps=None, epochs=10, clip=1.0, delta=1e-5):
    torch.manual_seed(0)
    model = nn.Sequential(nn.Linear(20, 64), nn.ReLU(), nn.Linear(64, 5))
    opt = torch.optim.SGD(model.parameters(), lr=0.1)
    dl, pe = loader, None
    if target_eps:
        pe = PrivacyEngine(accountant="rdp")
        model, opt, dl = pe.make_private_with_epsilon(
            module=model, optimizer=opt, data_loader=loader,
            target_epsilon=target_eps, target_delta=delta,
            epochs=epochs, max_grad_norm=clip)
    crit = nn.CrossEntropyLoss()
    for _ in range(epochs):
        for xb, yb in dl:
            if len(xb) == 0:      # Poisson sampling can yield empty batches
                continue
            opt.zero_grad()
            crit(model(xb), yb).backward()
            opt.step()
    model.eval()
    acc = (model(xte).argmax(1) == yte).float().mean().item()
    eps = pe.get_epsilon(delta) if pe else float("inf")
    sigma = opt.noise_multiplier if pe else 0.0
    return sigma, eps, acc

print("non-private", run())
for e in [0.5, 1, 3, 8]:
    s, eps, a = run(target_eps=e)
    print(f"target eps {e}: sigma={s:.3f} eps={eps:.2f} acc={a:.3f}")

The output on my machine was as follows. Treat it as illustrative: a different seed or hardware moves the third decimal.

Setting Noise multiplier Epsilon spent Test accuracy
Non-private 0 infinite 0.646
Target epsilon 0.5 4.06 0.49 0.641
Target epsilon 1 2.25 0.99 0.640
Target epsilon 3 1.09 2.99 0.638
Target epsilon 8 0.72 7.99 0.637

Two honest readings follow. First, the utility cost here is tiny, under one percentage point, which says more about the task than about DP. A linear-ish problem with ten thousand rows, twenty features and a small network is exactly where noise averages out across a large batch. Second, the price of tighter privacy shows up in the noise multiplier, which climbs from 0.72 to 4.06, and on harder tasks that is where accuracy disappears. The experiment also needs a caveat: the DP numbers include the effect of clipping with a fixed norm, and I did not tune the learning rate, clip norm or batch size per setting, so the table is a floor on what tuning could recover, not a benchmark of Opacus.

How Opacus is wired

PrivacyEngine.make_private wraps the module in a GradSampleModule, which computes per-example gradients using hooks on supported layers, replaces the optimizer with a DPOptimizer that clips and adds noise, and swaps the data loader for a Poisson-sampling one. The documented signature takes module, optimizer, data_loader, noise_multiplier and max_grad_norm, with poisson_sampling=True and clipping="flat" as defaults. After training, privacy_engine.get_epsilon(delta) returns the epsilon spent so far. Two practical gotchas follow from that architecture. Layers that mix examples, most famously batch normalization, break the per-example gradient assumption, so the usual fix is to replace them with group normalization. And per-sample gradients cost memory: a naive implementation stores a gradient per example per parameter, which is why large-model DP work relies on ghost clipping or similar tricks that never materialise the full per-example gradients. Opacus exposes alternative grad_sample_mode options for this purpose, and JAX makes per-example gradients cheap through vmap.

The hyperparameters that matter

Three knobs interact. The clip norm C should sit near the median gradient norm: set it too high and the noise, scaled by C, swamps the signal; too low and clipping biases every gradient toward a unit vector, which behaves like a normalised-gradient method with a changed objective. The noise multiplier sigma sets privacy per step. The batch size sets both q and the signal-to-noise ratio: noise standard deviation is sigma times C, but the signal is the sum of up to B clipped gradients, so larger batches improve signal-to-noise roughly linearly while increasing q only mildly. The consistent empirical finding in the literature is that DP training prefers larger batches, more careful learning-rate schedules and fewer epochs than ordinary training. I have not reproduced those studies here, so treat that as a reported pattern, not a result of the experiment above.

Attacks, Pretrained Models and LLM Fine-Tuning

Direct answer: Membership inference tests whether a specific record was in the training set. DP bounds any such attack: at a given false positive rate, the true positive rate cannot exceed e to the epsilon times that rate, plus delta. For language models, DP fine-tuning of a pretrained model is the practical route, because DP from scratch costs far more accuracy.

What epsilon means against a membership-inference attacker

The cleanest way to read epsilon is through the attacker. Suppose an adversary decides “in” or “out” for a target record. For an (epsilon, delta)-DP mechanism, any such test with false positive rate alpha has a true positive rate of at most e^epsilon times alpha plus delta. That bound is the operational definition of the privacy budget.

Run the numbers. At epsilon of 1, e^epsilon is 2.718, so an attacker who is wrong about non-members one time in a thousand can be right about members at most about 0.27 percent of the time (ignoring delta). At epsilon of 3, e^epsilon is about 20, giving a ceiling near 2 percent at the same false positive rate. At epsilon of 8, e^epsilon is about 2,981, so the ceiling at a 0.1 percent false positive rate is about 298 percent, a number above 1, which means the guarantee says nothing there. The bound only becomes informative again at false positive rates below roughly 0.03 percent. This is why very large epsilons are not merely weaker; for the strongest attack setting they are vacuous. It is also why practical attacks on models trained with epsilon of 8 or more can be far less successful than the worst-case bound permits. The bound is a ceiling, not a forecast.

Empirical audits complement the proof

The theoretical bound is an upper limit; the empirical question is how much leakage an actual attack achieves. A privacy audit trains models with and without canary records, runs the strongest available attack, and converts the observed true and false positive rates into an empirical lower bound on epsilon. If the empirical epsilon exceeds the claimed epsilon, there is a bug in the implementation or the accounting. The mismatch between sampling methods noted earlier is one source of such gaps. Auditing therefore plays the role of a unit test for your privacy claim. The Carlini et al. work cited above shows how to evaluate at low false positive rates, and an audit worth its name reports exactly that operating point.

DP fine-tuning of language models

Training a large language model from scratch with DP-SGD is currently very expensive in accuracy and compute: the noise is added to every parameter, and the noise norm grows with the number of parameters. The workable recipe, reported by Yu et al. (“Differentially Private Fine-tuning of Language Models”, 2021) and Li et al. (“Large Language Models Can Be Strong Differentially Private Learners”, 2021), starts from a model pretrained on public data and applies DP only to the fine-tuning step on private data. Two ingredients recur in those papers: very large batches, and either parameter-efficient methods such as LoRA that shrink the number of noised parameters, or memory-saving clipping techniques that make large batches affordable. Both papers report that fine-tuned models can approach non-private quality on several benchmarks at moderate epsilon; I have not reproduced the exact figures here and the claim depends on the task and the model.

There are two caveats a careful reader should carry. First, the DP guarantee covers only the private fine-tuning data; the pretraining corpus is treated as public, so anything it contains is outside the guarantee. Second, “record” has to be defined. For text, a record can be a sentence, a document or a user, and a guarantee at the sentence level says nothing about a user who contributed ten thousand sentences. User-level DP requires bounding each user’s contribution before clipping, which costs more budget.

Choosing Epsilon Honestly, and DP Against the Alternatives

Direct answer: There is no universally safe epsilon. Values below 1 give strong protection but cost real accuracy; values from 1 to 10 are the working range for most production ML; values above 10 give guarantees that are weak or vacuous against strong attackers. Report epsilon together with delta, the unit of privacy, the accountant and the sampling method.

How to pick a number you can defend

Start from the threat, not the library. Write down who the adversary is, what they already know, what harm a successful membership or reconstruction inference does, and what record means in your data (a row, a patient, a device, a user). Then choose the largest epsilon at which the bound in the previous section is still meaningful at the false positive rate you care about. For a clinical dataset where a single positive identification is a serious harm, that might be a low false positive rate such as 0.1 percent, which pushes the budget toward the low single digits. For a recommender trained on behavioural logs where the adversary’s gain is small, a larger epsilon may be an accepted business trade-off, provided it is labelled that way.

Three further rules keep the number honest. Include the whole pipeline: hyperparameter tuning on private data consumes budget, and so does every retraining on the same data, because the guarantees compose across runs. Prefer a final-model epsilon computed with the more precise accountant, and state it. And never compare epsilons across papers without checking that the neighbouring-dataset definition (add/remove versus replace-one), the unit of privacy and the delta are the same. An epsilon of 8 at record level and an epsilon of 8 at user level are different promises.

Group privacy and correlated data

DP protects a single record by default. If k records belong to the same person or are strongly correlated, the guarantee degrades to roughly k times epsilon for that group, with a growing delta. This is the standard group privacy property, and it is not a flaw in the maths, only a reminder that the definition matters. In IoT and digital twin settings, where one machine or one building produces millions of correlated sensor rows, the sensible unit of privacy is the asset or the tenant, not the row. A row-level epsilon of 1 on a telemetry stream with a slow-moving signal can protect much less than it appears.

Decision matrix: DP, federated learning, anonymization and synthetic data

Decision flow for differential privacy machine learning versus federated learning, anonymization and synthetic data

Figure 4: A decision flow for choosing a privacy technique. Where the raw data sits and whether a formal bound is required decide the branch.

These techniques solve different problems and are often stacked. The matrix compares them on the properties that decide a design.

Property DP-SGD or DP fine-tuning Federated learning alone Anonymization or redaction DP synthetic data
Formal leakage bound Yes, epsilon and delta No No Yes, if generated under DP
Raw data stays local No, one trusted trainer Yes No No for the generator step
Protects against membership inference Bounded by epsilon Not by itself Weak Bounded by epsilon
Utility cost Moderate to high, task dependent Low to moderate Low Moderate, and fidelity is limited
Reusable for many downstream tasks The model only The model only The dataset The dataset
Operational complexity Medium High Low Medium to high
Typical failure Epsilon too large to mean anything Updates leak, honest-but-curious server Re-identification by linkage Generator memorises records

Federated learning moves computation to the data, which matters for jurisdiction and bandwidth, but model updates can be inverted to recover training examples, so serious deployments add clipping and noise on the client or secure aggregation plus central DP. Anonymization and redaction are valuable for removing direct identifiers from text before it reaches a model, a theme covered in the sibling post on PII detection and redaction for LLM applications, but they carry no formal guarantee against linkage with outside data. Synthetic data is attractive because it can be shared, yet a generator trained without DP can memorise its training rows, and the question of quality control is treated in the sibling post on synthetic data for LLM fine-tuning and model collapse. A common, defensible stack is redact direct identifiers, then train or fine-tune with DP-SGD, then release only the model or only DP statistics.

Where the cost lands

The utility cost of DP is not uniform. Rare classes and rare tokens suffer most, because clipping and noise are indifferent to which examples are rare yet the signal for a rare example is weak to begin with. This is the fairness cost of DP: Bagdasaryan, Poursaeed and Shmatikov documented in “Differential Privacy Has Disparate Impact on Model Accuracy” (2019) that accuracy loss falls disproportionately on underrepresented groups. If your minority subgroup is the reason you built the model, measure per-group accuracy before and after, not just the average.

Trade-offs, Gotchas, and What Goes Wrong

DP is a guarantee about an algorithm, and implementations fail in ways that the proof does not cover. The most common failure is not the maths but the pipeline around it, and the list below is ordered roughly by how often it bites.

  • Accounting that does not match the sampler. The reported epsilon assumes Poisson sampling at rate q. Shuffling and slicing fixed-size batches, which is what most training loops do, changes the real privacy loss. Use the library’s own data loader wrapper.
  • Privacy leaks through hyperparameter search. Tuning learning rate, clip norm and architecture on the private data over twenty trials spends budget that the final run’s epsilon does not include. Tune on public proxy data, or account for the search explicitly.
  • Unsupported layers. Batch normalization and any layer that couples examples invalidate per-sample clipping. Opacus ships a validator that rejects such modules; read its message instead of silencing it.
  • Floating-point and RNG weaknesses. The textbook proof assumes ideal real-valued noise from a true random source. Opacus offers secure_mode with a cryptographically strong generator, which needs an extra dependency and slows training. Mironov showed in 2012 that naive floating-point implementations of the Laplace mechanism can break the guarantee, so this is not paranoia.
  • Large epsilon as a marketing label. An epsilon of 50 is mathematically differential privacy and operationally close to nothing. If the number cannot be defended against the attacker model, do not use the word as reassurance.
  • Public data contamination. DP fine-tuning guarantees nothing about the pretraining corpus, and if private records already appeared in the public data the guarantee is moot for them.
  • Side channels outside the model. Logs, caches, checkpoints saved mid-training before noise was applied, and evaluation sets built from private data are all outside the guarantee. The released artefact must be the only thing that leaves.
  • Compute and memory. Per-example gradients multiply memory pressure, and DP favours large batches, which pushes against GPU limits. Budget for a two to several times slowdown versus non-private training; the exact factor depends on the model and method, and I have not measured it here.

There is a harder conceptual limit as well. DP bounds what can be inferred about one record from the output, but it does not stop a model from learning facts about a population. If everyone with a given postcode has a given condition, a DP model will learn that, correctly, and that is a feature of the guarantee, not a violation. Teams that expect DP to hide statistical facts about groups are asking for a different property.

Practical Recommendations

Treat differential privacy as an engineering budget with an owner, not a checkbox. Begin by stating the privacy unit, the threat model and the target false positive rate for membership inference, then derive a candidate epsilon from the bound rather than adopting a number from a paper. Fix delta below the reciprocal of the dataset size.

Start from a pretrained model on public data whenever you can, and fine-tune with DP on the private data, using parameter-efficient adapters if the model is large. Use large batches, a clip norm near the median per-example gradient norm and as few epochs as the task allows. Run Opacus with the default Poisson data loader, replace batch normalization with group normalization, and call get_epsilon at the end of every run, logging epsilon, delta, accountant, sigma, q and steps next to the checkpoint.

Evaluate three things before release: utility on the full test set, utility per subgroup, and a membership-inference audit at low false positive rates. Compare the empirical epsilon with the claimed one. If you find a gap, find the bug before shipping.

  • Define the unit of privacy and document the neighbouring-dataset convention.
  • Pick epsilon from an attacker-centred bound; fix delta below 1/n.
  • Keep a single ledger of every training run that touches the private data.
  • Use Poisson sampling and the library’s accountant; name the accountant in reports.
  • Tune hyperparameters on public proxy data or account for the search.
  • Measure per-group accuracy and audit with a low-false-positive-rate membership test.
  • Release only the model; protect intermediate checkpoints and logs.
  • Combine with redaction and access control; DP is one layer, not the only one.

Frequently Asked Questions

What is a good epsilon value for differential privacy in machine learning?

There is no single safe value. Epsilon below 1 is considered strong and costs noticeable accuracy on hard tasks. Values between 1 and 10 are the common working range for deployed ML systems. Values above 10 offer little protection against strong membership-inference attackers because the bound e to the epsilon exceeds 1 at low false positive rates. Pick epsilon from your threat model, report delta and the unit of privacy, and verify with an empirical audit instead of relying on convention.

What is DP-SGD and how does it differ from regular SGD?

DP-SGD is stochastic gradient descent modified to satisfy differential privacy. It samples batches by Poisson sampling, computes a gradient for each example, clips each gradient to a fixed L2 norm, sums them, adds Gaussian noise proportional to that norm and then takes the usual optimizer step. Regular SGD averages unclipped gradients, so one outlier can dominate an update and leak information. DP-SGD also needs an accountant that tracks the cumulative privacy loss across steps.

Does differential privacy reduce model accuracy?

Usually yes, and the amount depends on the task, dataset size, model size and epsilon. Small, well-separated problems lose almost nothing, as in the experiment above where accuracy moved by less than one percentage point. Hard tasks, small datasets and rare classes lose more, and the loss is not evenly spread across groups. Pretraining on public data, large batches and parameter-efficient fine-tuning reduce the gap. Measure the cost on your own data, including per-group metrics.

What is the difference between epsilon and delta?

Epsilon bounds how much the probability of any output can change when one record is added or removed: the ratio stays within e to the epsilon. Delta is a small probability that this bound fails. Epsilon is the main privacy dial; delta should be much smaller than one over the number of records so that the failure case cannot expose individuals in practice. A pure epsilon-DP mechanism has delta equal to zero, like the Laplace mechanism, while Gaussian-noise mechanisms need a positive delta.

Is federated learning private by itself?

No. Federated learning keeps raw data on client devices, which reduces exposure, but the model updates sent to the server can leak information about training examples and can be attacked with inversion and membership inference. Federated learning becomes formally private only when combined with mechanisms such as clipped, noised updates and secure aggregation under a DP accountant. Treat FL as a data-locality tool and DP as the leakage bound, and use them together when both properties matter.

Can differential privacy prevent large language models from memorising training data?

DP fine-tuning bounds what any single training record contributes to the model, so verbatim extraction of that record is limited by epsilon. The guarantee applies only to the data covered by the DP step. Text memorised during public pretraining is outside it, and the definition of a record, such as a sentence versus a user, determines whose data is actually protected. Combine DP with deduplication, redaction of direct identifiers and extraction testing.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *