PII Detection and Redaction for LLM Applications: Presidio, NER and Reversible Tokenization

PII Detection and Redaction for LLM Applications: Presidio, NER and Reversible Tokenization

PII Redaction LLM Pipelines: Presidio, NER and Reversible Tokenization

Educational engineering analysis. Nothing here is legal advice; regulatory framing is a starting point for a conversation with your counsel and privacy officer.

The first time most teams discover that their chatbot is a PII firehose is not in a security review. It is when someone greps the trace store for a credit card pattern and gets forty thousand hits. Every prompt, every retrieved document chunk, every tool result and every model reply passes through places that were designed to be verbose, durable and searchable, and none of them were designed with personal data in mind. A working PII redaction LLM pipeline is the control that sits between your users and all of those stores.

It matters now because LLM applications have moved from demos to systems of record. They read support tickets, medical notes and bank statements, and they send that text to third-party model providers and to your own observability vendor. Redaction is cheap compared with a breach, but only if it is designed as a pipeline with measured recall, not bolted on as a regex.

This article maps where personal data leaks in LLM applications, lays out the detector families and how to combine them, shows a reversible tokenization vault, and gives runnable Microsoft Presidio code that I executed and timed.

What this covers: leak points, regex/NER/LLM detectors, Presidio analyzer and anonymizer, reversible pseudonymization, gateway versus SDK placement, streaming, multilingual limits, latency, GDPR and HIPAA framing, and a decision matrix.

Context and Background

Personally identifiable information (PII) is not a single data type. Regulators define it by identifiability, not by format. The GDPR defines personal data in Article 4(1) as any information relating to an identified or identifiable natural person, which includes a name but also an online identifier, and which can include a combination of innocuous fields that together single someone out. In the United States, HIPAA’s de-identification standard at 45 CFR 164.514(b) offers two routes: Expert Determination, where a qualified expert documents that re-identification risk is very small, and Safe Harbor, which removes 18 enumerated identifier categories and requires no actual knowledge that what remains could identify the individual. The HHS guidance describing both is linked in Further Reading.

That definitional gap is the first hard problem. A detector can find a nine-digit Social Security number with near-perfect precision, but it cannot know that “the pilot who landed on the Hudson in January 2009” identifies one specific person. Detection tooling finds identifier-shaped strings; privacy law cares about identifiability. Treat any redaction layer as risk reduction, not anonymization in the legal sense. Pseudonymized data, where a mapping to the original exists, remains personal data under GDPR, which is exactly the situation a reversible vault creates.

The second pressure is the OWASP Top 10 for LLM Applications. The 2025 list carries LLM02: Sensitive Information Disclosure, and its mitigations include sanitizing data so user data does not reach training, least-privilege access, and tokenization or redaction of sensitive content. That is the same control this article builds, and it overlaps with prompt-injection defences, since an injected instruction is often how data leaves. If you have not read the companion piece on agentic AI security and prompt injection, the exfiltration paths there are the adversarial version of the accidental leaks below.

On the tooling side, the open-source reference point is Microsoft Presidio, now maintained under the Data Privacy Stack organization. Its documentation describes an Analyzer that identifies PII in text, an Anonymizer that de-identifies it with configurable operators, an image redactor and a module for structured data. The project is explicit that automated detection cannot guarantee that all sensitive information is found, and recommends additional protections alongside it. I take that sentence as the design brief: layers, measured recall, and defence in depth.

Commercial alternatives exist from cloud providers and data-loss-prevention vendors, and traditional data loss prevention (DLP) products can be used at the egress proxy. They tend to be strong at structured identifiers and weak at conversational free text, which is the dominant shape of LLM traffic. The rest of this article assumes you want to own the pipeline, whichever detector you finally plug in.

Where PII Actually Leaks in an LLM Application

Before choosing a detector, enumerate the surfaces. Teams routinely protect the one they thought about (the prompt going to the provider) and miss four others.

PII redaction LLM leak points across prompts, RAG context, tool outputs, logs and fine-tuning data

Figure 1: Every path personal data takes through an LLM application. The gateway is the one choke point that sees all of them.

The figure shows inputs entering a gateway from three sources and fanning out to three sinks: the model provider, the log and trace pipeline, and, later, a training set built from those logs. Each deserves a sentence of its own.

Prompts carry whatever users type. Users paste emails, contracts and screenshots-transcribed-to-text, and they do it because the assistant is useful for exactly that. RAG context is worse in one respect: the user did not type it. A retrieval step can pull a chunk containing another customer’s details into the prompt, so redaction on user input alone does not cover it. Tool outputs from CRM lookups, database queries and web fetches enter the context as text the model reads and may repeat.

On the sink side, provider retention is governed by contract and by the provider’s data-handling terms, which you must read for each vendor and tier. Logs and traces are usually the largest exposure by volume, because observability tooling captures full request and response bodies by default and keeps them for weeks. Finally, fine-tuning and evaluation sets assembled from production traces silently inherit everything in those logs. Memorization of training data by language models is documented in the research literature, so once PII reaches a training set, deletion becomes a model problem, not a database problem.

The consequence is architectural. The control point must see inbound text from all sources, must run before the first sink, and must be the only code path to the provider. That is a gateway, discussed below.

The Reference Architecture: Detect, Resolve, Tokenize, Restore

Direct answer: A production PII redaction pipeline for LLM traffic runs four stages on every request: detect candidate entities with layered recognizers, resolve overlaps and apply per-entity thresholds, replace each span with a stable token stored in a per-session vault, and restore tokens in the model’s reply. Only tokens leave your trust boundary; the vault never does.

Think of the pipeline as a function with a measurable contract: given text, return spans with entity types and confidence, and then given the same text, return a transformed string such that no allow-listed entity type survives. Everything else in this section is about making that contract measurable and cheap.

Detector families and what each is good at

Detection splits into four families, and good systems combine them rather than choosing.

Regex plus checksum recognizers match structured identifiers. A credit card number is 12 to 19 digits, and the Luhn checksum rejects most random digit strings. An IBAN has a country-specific length and a mod-97 check. Presidio’s supported-entities documentation lists credit card and IBAN as pattern-plus-checksum recognizers, and email as pattern, context and RFC-822 validation. These are fast (microseconds), deterministic, explainable, and produce few false positives when a checksum applies.

Named entity recognition (NER) finds entities defined by language rather than format: people, organizations, locations. Presidio’s PERSON and LOCATION entities come from the NLP engine, which can be spaCy, Stanza or a transformers model. NER handles “Ask Priya to call me” where no pattern would help, but it costs milliseconds to tens of milliseconds per request on CPU, depends on casing and context, and degrades on short fragments, names that are also common words, and non-English text.

Context and deny lists boost or add detections using surrounding words and known values. The word “badge” next to a six-digit number raises the confidence that it is an employee ID. A deny list of your own customers’ account names or VIP names is cheap and high-precision, and often the highest-value recognizer a team adds.

LLM-based detectors ask a model to list sensitive spans. They catch contextual PII that no pattern or NER model flags, such as a rare disease mentioned in a support chat. They are the slowest and most expensive family, they are non-deterministic, and, uncomfortably, they require sending the unredacted text to a model. Use them on a local small model, or on asynchronous audit paths, not on the hot path to a third party.

Layered PII detection pipeline combining regex checksum recognizers, named entity recognition and deny lists

Figure 2: Detector layers run on normalized text, their candidates are merged, then overlaps are resolved against per-entity thresholds.

The ordering in Figure 2 matters. Normalization comes first because attackers and accidents both use Unicode tricks: full-width digits, zero-width joiners inside an email address, or “four one one one” spelled out. Normalize, or your regex never sees the string.

Recall, precision and why the thresholds are per entity

For redaction, the cost of the two errors is asymmetric. A false negative (missed PII) is a privacy incident. A false positive (over-redaction) degrades the answer quality, sometimes badly, if the redacted span was the thing the user asked about. Because the costs differ by entity type, one global threshold is the wrong design.

A reasonable policy sets a low threshold for entities where a miss is severe and false positives are cheap to tolerate (card numbers with a valid checksum, government identifiers), and a higher one for noisy entities such as PERSON from a small NER model or loosely patterned phone numbers. Section “Deeper Analysis” below shows real output where a six-digit fragment scores 0.01 as a driver licence and where the URL recognizer fires inside an email address. Without threshold and overlap logic, those pollute the output.

Measure on your own traffic. Public benchmark corpora rarely resemble your tickets, chat transcripts or clinical notes. Label a few hundred real, consented, access-controlled samples, compute recall per entity type, and track it on every detector change. I return to this in the recommendations.

Reversible tokenization versus irreversible redaction

Irreversible redaction replaces a span with a fixed label such as <PERSON> or deletes it. It is the safest and the right default for logs and training data, where nobody should ever see the original. Its weakness is functional: a reply that says “I have emailed ” is useless to the user, and two different people in one conversation both become <PERSON>, so the model cannot tell them apart.

Reversible tokenization (also called pseudonymization) replaces each distinct value with a stable, unique token and stores the mapping in a vault. The model reasons over <PERSON_e9c724> and <PERSON_91ab02> as separate entities, and the gateway swaps the originals back into the reply before it reaches the user. The model never sees the real values, yet the conversation stays coherent.

Two details decide whether this works. First, tokens must be stable within a session so the same email maps to the same token every time, which a keyed hash of the entity type and value gives you without storing counters. Second, token shape affects model behaviour. A token that looks like its type (<EMAIL_ADDRESS_b98096>) lets the model reason correctly about what it is; a token that is a random hex string invites the model to treat it as a code and sometimes to mangle it. Where downstream validators demand a format, use format-preserving tokens: a card-shaped token that passes the Luhn check but is not a real card, or a fake-but-valid-looking email at a reserved example domain. Format-preserving encryption schemes such as NIST SP 800-38G’s FF1 exist for this purpose; for LLM traffic a vault with generated surrogate values is usually simpler and avoids key-management coupling to the model’s context.

Sequence of a reversible tokenization gateway sending tokens to the LLM provider and restoring originals in the reply

Figure 3: The vault lives inside the trust boundary. The provider sees only tokens, and the gateway rehydrates the streamed reply.

Note one security property that is easy to miss. The vault turns a distributed leak (PII in many logs) into a concentrated one (PII in one store). That is deliberate and good, but it means the vault needs encryption at rest, per-tenant keys, a short time-to-live tied to the session, access auditing, and no export path to the analytics warehouse.

Gateway, SDK or Sidecar: Where Redaction Should Run

Placement is the decision teams revisit most, so it deserves a direct comparison rather than a default.

An SDK or middleware approach wraps each model call in application code. It is easy to adopt, has no extra network hop, and has access to rich context such as which field a string came from. Its weakness is coverage. Every service, notebook and agent framework must adopt it, one forgotten code path bypasses it, and upgrading the detector means redeploying every service.

A gateway is a proxy that all model traffic must traverse, with egress rules that block direct calls to provider domains. It gives one place to upgrade recognizers, one place to log redaction metrics, and a hard guarantee of coverage that is enforceable by network policy rather than by code review. Its costs are an extra hop, a new tier to operate and scale, and less semantic context: the gateway sees a JSON body, not the application’s knowledge that a field is a customer name.

A sidecar sits between the two. It runs next to each service, shares its lifecycle, and can be enforced by the service mesh, which keeps the hop local. It multiplies the number of detector instances (and NER model memory) by the number of pods, which is a real cost for spaCy or transformer models.

My recommendation for anything regulated is the gateway for enforcement, plus a thin SDK for hints: the SDK tags fields with known entity types (this string is a customer name), and the gateway uses those hints to skip detection and tokenize deterministically. Hints are cheaper and more accurate than inference, and the gateway remains the safety net for everything unhinted.

Redacting before the logs, not after

A subtle ordering bug is common. The gateway redacts the request to the provider but the tracing middleware, sitting earlier in the stack, has already serialized the original body into a span. Redaction must precede every sink, including your own telemetry. Practically that means the redaction step is the first thing the request touches after authentication, and trace exporters receive only the tokenized form. Exceptions and error messages deserve the same treatment, since stack traces frequently embed the failing payload.

Streaming responses

Server-sent event streaming breaks naive restoration. The model emits tokens a few characters at a time, so a vault token such as <PERSON_e9c724> can arrive split across chunks: <PERS then ON_e9c then 724>. A restorer that processes each chunk independently never matches the pattern and leaks the token (harmless) or, worse, passes a half-token to a client that renders it as markup.

The fix is a small holdback buffer. On each chunk, find the last opening delimiter; if no closing delimiter follows and the tail is shorter than the maximum token length, hold that tail until the next chunk. The code later in this article implements exactly that, and the added latency is bounded by the token length, a few tens of characters, not by the response length. Choose a token syntax with unambiguous delimiters to make this reliable; angle brackets are convenient but collide with HTML and XML content, so some teams prefer a rarer sentinel such as double curly braces or a private-use Unicode pair.

The inbound direction has a mirror-image problem for streaming inputs (voice transcripts, live chat). Entities straddle chunk boundaries there as well, so detection should run on a sliding window with overlap rather than on isolated fragments.

What about the model changing the token?

Models sometimes paraphrase. Given <PERSON_e9c724>, a model may write “Ms. e9c724” or lowercase the token. Three mitigations help. Instruct the model in the system prompt to copy placeholders verbatim. Prefer tokens that resemble natural entities of their type. And on the restore path, match case-insensitively and tolerate whitespace, while logging every orphan token that fails to resolve so you can measure the rate. I know of no published orphan-rate figure, so measure your own, and expect it to vary with model size and quantization.

Deeper Analysis: Presidio in Practice, With Runnable Code

The code below was executed on a two-vCPU Linux sandbox with presidio-analyzer and presidio-anonymizer 2.2.364 and spaCy’s small English model en_core_web_sm. Treat output scores as illustrative of behaviour, not as benchmarks, because they depend on the model and version.

Step 1: analyzer with a custom recognizer

Presidio’s AnalyzerEngine runs a registry of recognizers over text and returns RecognizerResult objects with entity type, span and score. The NLP engine is configured explicitly so that the model choice is visible in code rather than a download side effect. A custom recognizer takes a list of Pattern objects, each with a name, a regex and an initial score, plus optional context words that raise the score when they appear nearby.

from presidio_analyzer import AnalyzerEngine, Pattern, PatternRecognizer
from presidio_analyzer.nlp_engine import NlpEngineProvider

nlp = NlpEngineProvider(nlp_configuration={
    "nlp_engine_name": "spacy",
    "models": [{"lang_code": "en", "model_name": "en_core_web_sm"}],
}).create_engine()

employee_id = PatternRecognizer(
    supported_entity="EMPLOYEE_ID",
    patterns=[Pattern(name="emp_id", regex=r"\bEMP-\d{6}\b", score=0.6)],
    context=["employee", "badge", "staff"],
)

analyzer = AnalyzerEngine(nlp_engine=nlp, supported_languages=["en"])
analyzer.registry.add_recognizer(employee_id)

text = ("Hi, I'm Priya Raman (EMP-204518). Card 4111 1111 1111 1111, "
        "mail priya.raman@example.com, call +1 415-555-0132 from 203.0.113.9.")
for r in analyzer.analyze(text=text, language="en"):
    print(r.entity_type, round(r.score, 2), repr(text[r.start:r.end]))

Running this produced the following (card number is the standard test value, and the person is fictional):

CREDIT_CARD 1.0 '4111 1111 1111 1111'
EMAIL_ADDRESS 1.0 'priya.raman@example.com'
PERSON 0.85 'Priya Raman'
EMPLOYEE_ID 0.6 'EMP-204518'
IP_ADDRESS 0.6 '203.0.113.9'
URL 0.5 'example.com'
PHONE_NUMBER 0.4 '+1 415-555-0132'
US_DRIVER_LICENSE 0.01 '204518'

Read the output critically, because it is the argument for the whole resolution stage. The credit card and email score 1.0 because validation passed. PERSON scores 0.85 from the NER model. The phone number scores only 0.4, low enough that a naive threshold of 0.5 would let a real phone number through; that is why thresholds are per entity. The URL hit lies inside the email address, so it overlaps a higher-scoring span. And a driver-licence match at 0.01 sits inside the employee ID: a weak pattern firing on six digits.

Two of these are noise, one is a dangerous near-miss, and none of it is a bug. This is what recognizers do. The pipeline’s job is to turn it into a decision.

Step 2: resolve, tokenize and restore

The next listing adds the pieces Presidio leaves to you: per-entity thresholds with an allow-list, overlap resolution, a keyed-hash vault, and the streaming restorer. Entities not in the allow-list (here the URL and driver licence) are dropped by default, which makes the policy explicit and conservative about noise.

import hmac, hashlib, re
from dataclasses import dataclass, field

MIN_SCORE = {"PERSON": 0.6, "EMAIL_ADDRESS": 0.5, "CREDIT_CARD": 0.5,
             "PHONE_NUMBER": 0.4, "IP_ADDRESS": 0.5, "EMPLOYEE_ID": 0.5}

def resolve(results):
    """Keep allow-listed entities above threshold; on overlap keep the
    higher score, then the longer span."""
    keep = [r for r in results if r.score >= MIN_SCORE.get(r.entity_type, 1.1)]
    keep.sort(key=lambda r: (-r.score, -(r.end - r.start)))
    chosen = []
    for r in keep:
        if all(r.end <= c.start or r.start >= c.end for c in chosen):
            chosen.append(r)
    return sorted(chosen, key=lambda r: r.start)

@dataclass
class Vault:
    key: bytes
    fwd: dict = field(default_factory=dict)   # (type, value) -> token
    rev: dict = field(default_factory=dict)   # token -> value

    def token(self, etype, value):
        k = (etype, value)
        if k not in self.fwd:
            digest = hmac.new(self.key, f"{etype}|{value}".encode(),
                              hashlib.sha256).hexdigest()[:6]
            tok = f"<{etype}_{digest}>"
            self.fwd[k], self.rev[tok] = tok, value
        return self.fwd[k]

    def restore(self, text):
        return re.sub(r"<[A-Z_]+_[0-9a-f]{6}>",
                      lambda m: self.rev.get(m.group(0), m.group(0)), text)

def redact(analyzer, vault, text, language="en"):
    spans = resolve(analyzer.analyze(text=text, language=language))
    out, pos = [], 0
    for r in spans:
        out.append(text[pos:r.start])
        out.append(vault.token(r.entity_type, text[r.start:r.end]))
        pos = r.end
    out.append(text[pos:])
    return "".join(out)

class StreamRestorer:
    """Hold back a trailing partial token so split chunks restore correctly."""
    def __init__(self, vault): self.v, self.buf = vault, ""
    def feed(self, chunk):
        self.buf += chunk
        cut = self.buf.rfind("<")
        if cut != -1 and ">" not in self.buf[cut:] and len(self.buf) - cut < 40:
            ready, self.buf = self.buf[:cut], self.buf[cut:]
        else:
            ready, self.buf = self.buf, ""
        return self.v.restore(ready)
    def flush(self):
        out, self.buf = self.v.restore(self.buf), ""
        return out

On the sample sentence the redacted text came out as:

Hi, I'm <PERSON_e9c724> (<EMPLOYEE_ID_f50390>). Card <CREDIT_CARD_286a4b>, mail <EMAIL_ADDRESS_b98096>, call <PHONE_NUMBER_6b14b5>.

I then fed a reply containing two tokens through StreamRestorer in seven-character chunks, which split tokens mid-way, and confirmed the output equalled a one-shot restore. The URL and driver-licence noise never reached the output, and the IP address was handled in the first run but omitted from this shorter sentence.

A few design choices in this code are worth defending. The token digest is an HMAC with a per-session key, so token values cannot be precomputed across sessions or guessed by an attacker who knows the scheme, and the same value yields the same token within a session. The vault stores fwd and rev in memory here; in production they move to an encrypted store with a TTL. And the restorer returns unknown tokens unchanged, which is the safe failure mode: an unresolved token is a cosmetic defect, not a leak.

Presidio’s own anonymizer and its limits

Presidio ships an Anonymizer with built-in operators: replace (with new_value, defaulting to the entity type label), redact, mask (chars_to_mask, masking_char, from_end), hash (sha256 by default, sha512 optional, with a salt), encrypt with a key, and custom taking a lambda. Operators are configured per entity as OperatorConfig("replace", {...}), and a "DEFAULT" key covers entities without a specific entry. Only encrypt is reversible, through DeanonymizeEngine with a matching decrypt operator.

Use the built-ins for the irreversible paths: replace or redact for logs and training data, mask for display. For the LLM round trip I prefer a custom vault over encrypt for two reasons. Ciphertext is long and high-entropy, which inflates token counts and invites the model to corrupt it. And encryption makes the key the recovery mechanism, so every service that can decrypt is a service that can leak; a vault with a lookup and a TTL centralizes that authority. Encryption is the right tool when you need stateless restoration across services, such as passing redacted records through a queue that a later worker restores.

Measured latency, labelled honestly

The numbers below are single-process medians from my sandbox with the small spaCy model, run after a warm-up call, using the sample sentence repeated to vary input size. They are illustrative of the shape of the curve, not a benchmark of your hardware or of larger models.

Input size Characters Median redact time
1 sentence 112 about 14 ms
8 sentences 896 about 46 ms
32 sentences 3,584 about 177 ms

Two lessons fall out. Cost scales with text length because the NER pass dominates, so detection time is roughly linear in the characters scanned; a RAG prompt with ten retrieved chunks costs far more than the user’s question. And a transformer-based NER model on CPU will be materially slower than the small spaCy model used here, while a GPU changes the picture again, so the budget conversation must be run against your actual model.

Practical tactics follow. Cache redacted RAG chunks at index time, so retrieval returns text already tokenized, and keep a chunk-level vault so tokens are stable across queries. Run regex and checksum recognizers first and skip NER on obviously structured payloads. Batch short strings into one analyzer call. And set a latency budget (for example, redaction must stay under a fixed fraction of the end-to-end time to first token) with a circuit breaker whose failure mode is fail closed for regulated routes: if redaction times out, the request is rejected rather than sent raw.

Decision flow for choosing redaction placement and detector depth by data class, with a recall gate

Figure 4: Route data by class, pick detector depth, and gate release on measured recall rather than on hope.

Multilingual and domain limits

Presidio’s predefined pattern recognizers are largely English and US centric; its entity list includes country-specific identifiers, but coverage varies by country and each is a pattern, not a guarantee. The NER side is configured per language, and its quality tracks the quality of the underlying model for that language. Names in languages without capitalization conventions, transliterated names, mixed-script text and code-switched chat all reduce recall, and a small English model run over Hindi or Arabic text will silently miss most people. If your users write in many languages, run language identification first, route to a model per language, and measure recall per language. Do not assume the English number transfers.

Domain shift hurts as well. Clinical notes use abbreviations and eponyms; chat logs use handles and nicknames; source code and stack traces contain tokens that look like keys. Secrets such as API keys and passwords are a separate class from PII and deserve their own recognizers, with entropy checks and provider-specific prefixes, because a leaked credential is often worse than a leaked name.

Two points are useful for engineers. Under the GDPR, pseudonymised data is still personal data, so a reversible vault reduces risk but does not take the system out of scope. Article 25 requires data protection by design and by default, and a redaction gateway is a concrete, auditable instance of that principle. Under HIPAA, Safe Harbor’s 18 identifier categories include more than names and numbers: dates more specific than a year, geographic units smaller than a state, and any other unique identifying code. A detector tuned for the “obvious” entities will not satisfy that list, which is why Expert Determination is often used for text-heavy data. Whether a given pipeline meets either bar is a legal and statistical question; the engineering contribution is evidence, namely measured recall per entity, retained samples of misses, and audit logs.

Trade-offs, Gotchas, and What Goes Wrong

The honest limit is that no detector reaches perfect recall, and the long tail is where the incidents are. Rare names, unusual formats, free-text descriptions that identify someone indirectly, and adversarial obfuscation will get through. Defence in depth means you assume a miss will happen and bound its impact: short log retention, access controls on traces, provider zero-retention terms where available, and egress rules that stop raw traffic.

Over-redaction breaks answers. Replace every location and the assistant cannot answer a question about the nearest branch. Replace a person’s name inside a quoted legal document and summarization quality falls. The remedy is route-specific policies and, for quality-sensitive tasks, reversible tokens so the model keeps entity identity without values.

Tokens leak structure. A token that preserves length, format or the last four digits leaks information by design. Preserve only what the downstream consumer needs. Similarly, deterministic tokens across sessions let anyone with log access correlate users, so scope keys per session or per tenant.

The vault is a crown jewel. It concentrates everything you worked to protect. It needs its own threat model, including insider access, backup copies and the debug endpoints people add during incidents.

Prompt injection can ask for the vault. A malicious document can instruct the model to output every token it has seen, or to emit a token inside a link whose rehydration lands in an attacker-visible URL. Restore only in the user-facing response path, never inside tool-call arguments sent to external systems, and never rehydrate tokens inside URLs or markdown image links the model produced.

Embeddings are not anonymous. If you embed raw text for retrieval, the vectors and the stored chunks carry PII and may be inverted to approximate the original text. Embed the tokenized text, and treat the vector store as in scope for retention and deletion requests.

Evaluation drifts. Models, spaCy versions and regex edits change recall silently. Pin versions, keep a regression suite of labelled samples, and run it in CI so a dependency bump cannot lower recall unnoticed. The same discipline applies to the supply chain of the detector models themselves, which is the subject of the piece on AI model supply chain security and provenance.

The last gotcha is organizational. A redaction layer creates a false sense of safety, and teams then relax other controls. Keep it labelled as one control among several. The breach analysis in the Bitget incident write-up is a reminder that third-party and zero-day exposure routes data around the controls you were proudest of.

Practical Recommendations

Start from data classes, not tools. Classify each route of your application as regulated, internal or public, and assign detector depth accordingly: regulated routes get the full layered pipeline, a gateway, fail-closed behaviour and a vault; internal routes get regex plus NER with irreversible redaction in logs; public content gets log scrubbing only.

Build the evaluation set before you tune anything. Collect a few hundred consented, access-controlled real samples per major language and route, label entities, and report recall and precision per entity type. Gate releases on recall for the severe entities, and publish the numbers internally so decisions about thresholds are visible. Add every miss from production review back into the set.

Prefer irreversible redaction wherever the original is not needed, and reserve the vault for the narrow set of flows where the user must see real values in the reply. Keep vault lifetimes to the session, key them per tenant, encrypt at rest, and audit reads.

Finally, put redaction in front of every sink, including your own telemetry, and use egress policy to make the gateway the only path to model providers.

  • Classify routes and map each to a detector depth and a failure mode.
  • Run layered recognizers: checksum regex, NER, deny lists, and secrets scanning.
  • Set per-entity thresholds and an explicit allow-list; drop everything else.
  • Resolve overlaps deterministically and log which recognizer won.
  • Redact before traces, logs, caches and embeddings, not just before the provider.
  • Use keyed, session-scoped tokens; restore only in the user-facing path.
  • Handle streaming with a holdback buffer and log orphan tokens.
  • Pin detector and model versions and run a recall regression in CI.
  • Measure latency per input size and set a fail-closed timeout for regulated routes.

Decision matrix

Approach Strengths Weaknesses Best fit
Regex plus checksum only Microseconds, explainable, few false positives for structured IDs Misses names, places and free-text PII Cards, IBANs, emails, secrets in logs
Presidio with NER Open source, extensible, covers names and locations, reversible via your vault English-centric defaults, tens of ms on CPU, needs tuning General LLM gateway, mixed content
Transformer NER or GLiNER-style models Better contextual recall, multilingual options Heavier compute, GPU often desirable, versions to pin High-recall regulated text
LLM-based detector (local) Catches contextual and indirect PII Slow, non-deterministic, costly Async audits, second-pass review
Managed cloud PII service or DLP Low operations burden, vendor-maintained models Data leaves your boundary, per-call cost, less control Teams without ML platform capacity

Frequently Asked Questions

What is PII redaction for LLM applications?

It is a control that finds personal data in text headed for, or returning from, a language model and replaces it with labels or tokens before the text reaches the provider, logs or training sets. It usually runs in a gateway and combines regex, checksum validation and named entity recognition. The goal is to reduce exposure and meet data-protection obligations, not to guarantee that nothing sensitive ever passes.

Does Microsoft Presidio catch all PII?

No. Presidio’s own documentation states that automated detection cannot guarantee that all sensitive information is found and recommends additional protections. It performs well on structured identifiers with checksums, and moderately on names and places depending on the NLP model. Recall varies by language and domain, so measure it on labelled samples from your own traffic and layer extra recognizers and controls on top.

What is the difference between redaction, masking and reversible tokenization?

Redaction removes or replaces a value with a fixed label and cannot be undone. Masking hides part of a value, such as all but the last four digits. Reversible tokenization swaps each distinct value for a unique token and stores the mapping in a vault so the original can be restored in the response. Reversible approaches keep conversations coherent but count as pseudonymization, which remains personal data under GDPR.

Should I redact at an API gateway or inside the application?

For regulated data, enforce at a gateway, because network egress rules can guarantee that no service bypasses it and the detector upgrades in one place. Use an SDK only to attach hints, such as which fields are customer names, so detection is cheaper and more accurate. Sidecars suit service-mesh environments but multiply the memory cost of NER models across every pod.

How do I restore tokens when the response is streamed?

Buffer the tail of the stream whenever it ends with an opening delimiter that has no closing one yet, and release it when the next chunk completes the token. This keeps latency bounded by token length. Restore only in the user-facing path, log tokens that fail to resolve, and never rehydrate tokens inside URLs, links or tool-call arguments sent to external systems.

How much latency does PII detection add?

It depends on text length, hardware and the NER model. In my sandbox run with Presidio and spaCy’s small English model on two vCPUs, a single sentence took about 14 ms and a 3,500-character input about 177 ms; these are illustrative, not benchmarks. Larger transformer models cost more on CPU. Cache redacted RAG chunks at index time and set a fail-closed timeout.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *