OpenAI Distillation Campaign: How CoT Extraction Works and How to Defend
On October 1, 2026, OpenAI described a coordinated effort to pull the hidden reasoning out of its models at scale, and said it had disrupted it. According to press coverage of the announcement, the activity ran through July, peaked on July 24 and 25 at roughly 16,000 requests from more than 4,000 users, and was shut down by July 28. A central cluster was tied to people associated with Moonshot AI, the company behind the Kimi models. Moonshot has not, in the coverage I could access, responded to the specific allegation. The technique has a name that is becoming a staple of AI security vocabulary: the model distillation attack.
It matters now because the economics are lopsided. Frontier labs spend enormous sums producing a model that reasons well, and a rival who can harvest enough of that reasoning can train a cheaper imitation for a small fraction of the cost. For anyone running an LLM API, the same playbook is aimed at you the moment your model is worth copying.
You will leave with a working mental model of how distillation works, why chain-of-thought traces are the prize, what OpenAI has said it did, and a concrete detection and mitigation stack you can apply to your own endpoints.
What this covers: the mechanics of distillation and its legitimate uses, what is publicly known about the OpenAI incident (with attribution), the economics and legal terrain, detection signals, defensive architecture, trade-offs, and a practical checklist.
Context and Background
Knowledge distillation is an old, respectable idea. Geoffrey Hinton, Oriol Vinyals and Jeff Dean formalized it in the 2015 paper “Distilling the Knowledge in a Neural Network” (arXiv:1503.02531): train a small student to match the output distribution of a large teacher, and the student inherits far more than hard labels could teach it. Everyday practice has drifted from that original formulation. Most modern LLM distillation does not match logits, because API customers never see them. It does something simpler: it generates text with the teacher and fine-tunes the student on that text, a technique often called sequence-level distillation or, informally, “training on synthetic data from a stronger model.”
That simplicity is why it is hard to police. A fine-tuning set built from teacher outputs looks, from the outside, like any other dataset. The only place the act is visible is at the teacher’s API, as a pattern of queries.
Distillation is not inherently illegitimate. Labs distill their own large models into smaller ones constantly, and many providers sell distillation as a product feature. Our own comparison of LoRA, QLoRA, full fine-tuning and distillation for small language models covers the legitimate version, where you own the teacher or have a licence. The problem is the adversarial version, where the teacher belongs to someone else and the terms of service forbid using its outputs to build a competing model.
What changed with reasoning models is the value of what leaks. A classic chat completion yields a final answer. A reasoning model additionally produces an internal chain of thought: a long scratchpad of decomposition, false starts, checks and corrections. That trace is the training signal that teaches a student how to reach answers, not just what they are. Providers responded by hiding it. OpenAI has exposed only summaries of reasoning for its o-series and later models, and in some API configurations returns reasoning as an encrypted item that the client can pass back but not read (see OpenAI’s reasoning-model documentation at platform.openai.com/docs). Hiding the trace turns it into a target.
The incident also did not arrive in isolation. The reporting I reviewed says Google, Anthropic and US officials have previously made similar accusations against Chinese AI developers, and one outlet reports that Anthropic separately said Moonshot-linked actors relayed nearly 300,000 customer requests through 5,380 fraudulent accounts over ten days. I could not verify that Anthropic figure against a primary source, so treat it as reported. The pattern across labs is consistent enough to plan around, whatever the final facts of any single case.
What Is a Model Distillation Attack and How Does Chain-of-Thought Extraction Work?
A model distillation attack is the use of a victim model’s API outputs, at scale and against its terms of service, to train a competing model. Chain-of-thought extraction is the high-value variant: the attacker tries to recover the model’s hidden reasoning, not just its answers, because reasoning traces transfer capability much more efficiently than final answers alone.

Figure 1: The basic model distillation attack pipeline. Every stage except the first is invisible to the teacher; only the bulk query stage leaves a footprint at the API.
The figure compresses a pipeline that is mundane in each step. The attacker builds a prompt set aimed at the capabilities they want, typically hard maths, code, science and multi-step planning. They send it through many accounts to stay under per-account limits. They collect responses, filter for correctness using verifiers or majority agreement, and fine-tune a student on the survivors. Only the query stage touches the victim, which is why every effective defense operates there or on what the response contains.
Why reasoning traces are the prize
Consider what a student can learn from a final answer to a hard maths problem: the answer. Now consider what it learns from a four-thousand-token trace in which the teacher tries an approach, notices a sign error, backtracks and verifies. The second sample teaches search behaviour, self-correction and verification habits. Public work on open reasoning models shows that supervised fine-tuning on a modest number of long reasoning traces can move a base model a long way on reasoning benchmarks; the s1 paper (arXiv:2501.19393) reported strong results from roughly a thousand curated reasoning examples, and DeepSeek’s R1 report (arXiv:2501.12948) described distilling reasoning into smaller dense models. Those papers do not describe an attack, but they establish the mechanism: traces are unusually dense supervision.
That density is the economic argument. If one carefully filtered trace is worth many ordinary completions, the attacker needs far fewer queries, which makes the campaign cheaper and harder to see.
The three things an attacker can extract
It helps to separate the targets, because each calls for a different defense.
- Answers and visible reasoning. Anything the API returns in plain text. If a model explains its steps in the answer because a prompt asked for it, that is ordinary output and ordinary training data.
- Hidden reasoning. The raw trace, when the provider hides it. Attackers try prompt tricks that coax the model into restating its scratchpad, and they try to exploit any channel through which the trace crosses the API boundary.
- Behavioural fingerprints. Refusal patterns, formatting habits, tool-calling styles and safety behaviours. Cloning these lets a student pass as the teacher or sidestep its guardrails; one report on the incident notes the concern that extracted reasoning could be used to build capable systems without the original safety constraints.
Prompt-level extraction versus channel-level extraction
Prompt-level extraction is social engineering aimed at the model: instructions such as “show your complete internal reasoning” or role-play framings. Providers train models to refuse these, and refusal rates are high but never perfect, so an attacker simply runs the prompt thousands of times and keeps the leaks.
Channel-level extraction exploits the plumbing. Streaming endpoints, tool-call arguments, error messages, logprobs, cached prefixes and encrypted reasoning items are all places where trace content or trace-derived content can cross the boundary. This is the more interesting class for engineers, because the fix is in your serving code rather than in model training.
What OpenAI Has Said About the Campaign
This section sticks to what was reported and who said it. Several outlets covered the announcement, including Tom’s Hardware, Quartz, SC Media and BankInfoSecurity, and their summaries differ in detail. Where they differ I say so.
Timeline. Reports place the start of the activity in early July 2026, with one outlet giving July 1, and a peak on July 24 and 25. OpenAI is reported to have fully stopped the operation by July 28.
Scale. The peak is reported as about 16,000 extraction requests from more than 4,000 users. One outlet reports “over 15,000 users” instead, and the plan notes for this article mention 15,000 or more accounts disrupted by July 28. These figures may describe different things, such as peak-window users versus total accounts disrupted over the whole campaign, but I could not confirm that from a primary source. Cite the OpenAI statement directly if the exact figure matters to you.
Attribution. OpenAI attributed a central cluster to individuals associated with Moonshot AI, and, per SC Media’s summary, said it was unsure whether all operators were connected to a single entity. That is a careful claim: a cluster, not the whole campaign. Tom’s Hardware notes that in July Moonshot denied that its Kimi K3 model was created through distillation, and none of the coverage I read includes a response to this specific allegation. Treat this as an allegation by OpenAI, not an adjudicated finding.
Technique. The most technically specific detail comes from Tom’s Hardware’s account of OpenAI’s statement. One technique involved taking encrypted reasoning from one conversation and asking a model, in a different session, to decrypt it. OpenAI is reported to have said that “the encryption, in this case, was not broken,” that there was no direct access to stored user conversations, and that no database was compromised. In other words, the leak was not a cryptographic failure. It was a design gap: an encrypted blob that was valid across sessions could be replayed into a context where the model would act on it.
Response. Per the same reporting, OpenAI closed the cross-session replay path, added detection for streamed output that exposes reasoning, strengthened protections across users, workspaces, organizations and model families, disrupted accounts including those reaching it through third-party providers, and shared intelligence with other AI companies and government programmes.
An adjacent research result. Tom’s Hardware also points to a paper titled “Stealing Reasoning Traces from Proprietary LLM APIs,” dated August 10, which reportedly showed the vulnerability class across OpenAI, Anthropic and Google, and reports that after the providers acknowledged it the researchers could not repeat the attacks. I have not read the paper itself, so I am relaying the press description only.
The Economics and the Legal Terrain
Understanding why the attack is attractive explains why detection has to be economic as well as technical.
A back-of-envelope cost model
The numbers below are illustrative, not measured, and exist to show the shape of the trade. Suppose an attacker wants 200,000 high-quality reasoning samples. Suppose each sample consumes 1,500 input tokens and 6,000 output tokens, and suppose output tokens cost around $10 per million (a plausible order of magnitude for a frontier reasoning tier, though prices vary and change). The output bill is 200,000 × 6,000 = 1.2 billion tokens, or roughly $12,000. Add input and a safety margin and the total is in the low tens of thousands of dollars.
Compare that with the cost of producing the same capability from scratch: pre-training compute, data licensing, and a large reinforcement-learning programme for reasoning. Public estimates for frontier training runs are in the tens to hundreds of millions of dollars, though labs rarely disclose exact numbers. Even if distillation recovers only a fraction of the capability, the ratio is hugely favourable to the copier. That asymmetry is the root cause. Defenses cannot make extraction impossible; they can only raise its price and cut the quality of what is recovered.
Why attackers spread across thousands of accounts
If each account can safely send only a few hundred requests before tripping limits, 200,000 samples need hundreds of accounts. A campaign with 4,000 users, as reported, fits this logic: the attacker trades account-creation cost for stealth. The reported use of third-party providers to reach the API adds another layer, since resellers and gateways dilute the identity signal OpenAI sees. Anthropic’s reported figure of 5,380 fraudulent accounts relaying nearly 300,000 requests over ten days, if accurate, works out to roughly 55 requests per account. Each account then looks like a light user.
Terms of service, and what law adds
The primary instrument is contractual. Major providers prohibit using outputs to develop competing models; OpenAI’s usage terms contain such a clause, and you should read the current text on the OpenAI site rather than rely on my paraphrase. Breach of contract allows account termination and, in principle, damages, but enforcement across borders is hard and slow.
Copyright offers a weak lever. Model outputs generated by a machine have uncertain protection in many jurisdictions, and the training-data fights that rage around frontier labs cut both ways, since those labs themselves rely on fair-use arguments for their own training. Trade-secret law is a better fit for hidden reasoning that the provider takes reasonable steps to protect, and unauthorized access theories such as the US Computer Fraud and Abuse Act are sometimes floated for deliberate circumvention of access controls, though applying them to API misuse is legally untested territory. I am not a lawyer and none of this is legal advice; the practical takeaway is that technical controls and account enforcement do most of the work, with legal and policy channels (export-control debates, government intelligence sharing, as the reporting mentions) playing a supporting role.
The safety angle
Labs also frame distillation as a safety issue. A student trained on a teacher’s capabilities but not its alignment training may reproduce the first without the second. This argument is contested, since open-weight models can also be fine-tuned for safety, but it is part of why providers share intelligence with government programmes and with each other. For defenders the practical implication is that detection signals are worth sharing across organizations: the same actors tend to probe several providers.
Reference Architecture: Defending an LLM API Against Distillation
Layered defense beats any single control. The architecture below sits between untrusted clients and your model servers, and has four jobs: decide how much reasoning leaves the building, make accounts expensive to fake, score behaviour, and close channel leaks.
Layer 1: Decide your trace exposure policy
The most powerful lever is how much of the model’s reasoning you return at all.

Figure 2: Four trace exposure policies, ordered from most to least leaky. Encrypted traces look safe but carry a replay risk unless bound to their context.
Returning the full raw trace maximizes developer debuggability and trust, and it also hands attackers a perfect training set. This is why several providers moved away from it. Returning only a summary is lossy: the summary is a second model’s paraphrase, which drops the false starts and corrections that carry the training signal. It is still some signal, and attackers can train on summaries, but the sample efficiency is lower. Hiding the trace entirely gives the least leakage, at the price of debuggability, and the model’s answer still carries some reasoning when it explains itself.
The encrypted-trace option exists because multi-turn reasoning and tool use work better when the model can see its own earlier reasoning. The provider returns an opaque blob and the client sends it back unchanged on the next turn. The reported OpenAI incident shows the failure mode in this design: if that blob is not cryptographically bound to the session, user, workspace and model that produced it, anyone who obtains it can submit it somewhere else and try to coax a model into revealing it. The fix is a binding. Include the identifiers of the originating context in the authenticated data of an AEAD (authenticated encryption with associated data) scheme such as AES-GCM, so decryption fails, or the blob is simply discarded, when presented elsewhere. OpenAI’s reported fix of strengthening protections across users, workspaces, organizations and model families reads like this idea applied broadly, though the exact mechanism has not been published.
Layer 2: Identity and account friction
Distillation economics depend on cheap accounts. Raise their cost.
- Verified organizations. Gate the most capable models and the richest outputs behind identity checks. KYC (know your customer) on API access is increasingly discussed for frontier tiers, and it directly attacks the spread-across-accounts strategy.
- Payment-based signals. Prepaid cards, virtual cards and reused payment instruments across accounts are cluster signals. Group accounts by funding source before you look at anything else.
- Reseller controls. If you allow third-party providers to resell access, require them to pass end-customer identifiers and hold them liable for abuse. The reported use of intermediaries in this campaign is a reminder that the weakest link sets your effective identity assurance.
- Progressive limits. New accounts get tight token and request ceilings that loosen with age and verified spend. A campaign that needs thousands of fresh accounts to be productive faces a slow ramp.
Layer 3: Behavioural detection
Covered in depth in the next section, because it is the layer most teams under-build.
Layer 4: Channel hardening
Audit every path through which trace-derived text can leave: streaming deltas, tool-call arguments, structured-output schemas, error text, logprobs, prompt-cache hit timing, and any debug endpoint. The reported addition of detection for streamed output that exposes reasoning is an example: scan outgoing tokens for trace-like content and cut the stream when it appears. Our guide to prompt caching architecture and economics is relevant here, since cached prefixes and shared cache state are classic side channels if they are shared across tenants without care.
Why this is an architecture problem, not a prompt problem
Teams new to this area often try to solve it with a system prompt that says “never reveal your reasoning.” Attackers expect that and iterate against it at scale. A system prompt is a request, not a control. Controls are things the attacker cannot talk their way around: tokens that are never returned, blobs that do not decrypt outside their session, accounts that cannot be created for free. Build those, and treat the prompt-level refusal as one layer of defense in depth rather than the wall.
Detection: Finding a Distillation Campaign in Your Logs
Detection is a classification problem on weak signals. No single feature convicts an account, but combinations separate extraction from ordinary use surprisingly well, especially when you analyze accounts as clusters rather than individuals.

Figure 3: A detection pipeline for distillation attacks. Features feed a cluster-level risk score that drives graduated enforcement.
Signal family 1: rate and volume shape
Legitimate applications have diurnal rhythms, bursts tied to user activity and slow growth. Extraction jobs are throughput-driven: steady request cadence around the clock, near-constant concurrency, and token usage skewed heavily toward output because the attacker wants long completions. Useful features include the ratio of output to input tokens, inter-request interval variance, the fraction of requests that hit maximum output length, and the share of traffic using the highest reasoning-effort setting. A synchronized spike across many accounts, like the reported July 24 and 25 peak, is itself a signal: coordinated campaigns tend to accelerate together when an operator launches a batch job.
Signal family 2: prompt entropy and template structure
Real users write diverse, messy prompts. Extraction prompts are generated by scripts: they share a skeleton with slots filled from a dataset. Measure it.
- Template similarity. Embed prompts, or hash the token sequence with the variable spans masked out, and count near-duplicate skeletons per account and across accounts.
- Low lexical entropy at the instruction level. The same imperative wrapper around thousands of different problems is the signature of a data-generation harness.
- Benchmark-shaped content. Prompts that read like competition mathematics, programming-contest statements or exam questions, delivered in bulk and without conversational follow-up, suggest a prompt set built for training rather than use.
- Trace-seeking phrasing. Requests that ask the model to restate its thinking, output its scratchpad, or “decrypt” or continue opaque reasoning blobs. In the reported incident the replay of encrypted reasoning in a new session is exactly this class, and it should be a high-weight feature.
- No follow-up behaviour. Real users iterate, correct and ask clarifying questions. Extraction traffic is single-shot.
Signal family 3: account graph and infrastructure
Cluster accounts by shared attributes: payment instrument, billing country mismatch, IP ranges and ASNs (autonomous system numbers), TLS and HTTP client fingerprints, API key creation timing, organization naming patterns, and referring reseller. Graph connected components often reveal a campaign that is invisible at the account level. A group of 4,000 accounts that each look unremarkable collapses into a few dozen clusters once you link on funding and infrastructure. This is where intelligence sharing helps, because a provider can match indicators contributed by peers.
Combining signals into a score
A practical approach is a gradient-boosted model or even a hand-weighted logistic score over the features above, evaluated per account and per cluster. Train on past confirmed abuse, and keep a labelled hold-out of benign heavy users, such as evaluation harnesses and data-labelling vendors, because those are your main false-positive source. I am not aware of a published benchmark for distillation detectors, so any precision and recall numbers you see quoted should be treated sceptically and measured against your own traffic.
Graduated enforcement matters. Low scores pass. Medium scores trigger throttling, forced re-verification or a reduction in reasoning-effort tier, none of which breaks a legitimate customer permanently. High scores trigger suspension of the cluster, revocation of keys and, where appropriate, sharing of indicators. Disrupting the cluster rather than one account at a time is what prevents an operator from simply rotating to the next login.
Honeypots and canaries
Two techniques deserve more attention than they get. First, canary prompts: plant distinctive, unlikely question-and-answer pairs into responses to flagged traffic, then watch whether competitor models later reproduce them. This is a form of data watermarking and gives evidence of use after the fact. Second, output watermarking at the token level, where the sampler subtly biases token choices according to a keyed pattern so that text can later be statistically identified as coming from your model. Research on LLM watermarking exists (for example Kirchenbauer et al., arXiv:2301.10226), but robustness against paraphrasing and mixed corpora is limited. Treat watermarks as supporting evidence for enforcement and for disputes, not as a standalone defense.
Putting It Together: Request Flow and a Working Sketch
The pieces above need to compose into a single request path with low latency overhead.

Figure 4: Request flow for a protected reasoning endpoint. Scoring is inline and cheap; trace stripping is mandatory; analyst review is asynchronous.
The gateway authenticates the key, extracts lightweight metadata and prompt features, and asks the detector for a score. The detector must answer in a few milliseconds, so it relies on precomputed account and cluster risk plus cheap per-request features, with the heavy clustering running offline. The model’s response then passes through trace policy enforcement before leaving: raw reasoning is removed or summarized, any encrypted item is bound to its context, and the outgoing stream is scanned. Escalations flow asynchronously to analysts, whose decisions feed back as revocations and tighter limits.
A minimal scoring sketch
The code below is a teaching sketch of feature extraction and a hand-weighted score. The weights are placeholders, not tuned values, and a production system would learn them from labelled data.
import math
from collections import Counter
def shingle_hash(prompt: str, mask_numbers: bool = True) -> int:
"""Hash a prompt skeleton so near-identical templates collide."""
toks = prompt.lower().split()
if mask_numbers:
toks = ["<n>" if t.strip(".,").isdigit() else t for t in toks]
return hash(" ".join(toks[:12])) # first 12 tokens approximate the wrapper
def entropy(counts: Counter) -> float:
total = sum(counts.values())
return -sum((c / total) * math.log2(c / total) for c in counts.values())
def account_features(reqs: list[dict]) -> dict:
"""reqs: [{prompt, in_tok, out_tok, ts, effort}] for one account window."""
n = len(reqs)
skeletons = Counter(shingle_hash(r["prompt"]) for r in reqs)
top_skeleton_share = skeletons.most_common(1)[0][1] / n
out_in = sum(r["out_tok"] for r in reqs) / max(1, sum(r["in_tok"] for r in reqs))
high_effort = sum(r["effort"] == "high" for r in reqs) / n
trace_seek = sum(
any(k in r["prompt"].lower() for k in
("your reasoning", "scratchpad", "decrypt", "chain of thought"))
for r in reqs) / n
gaps = [b["ts"] - a["ts"] for a, b in zip(reqs, reqs[1:])]
mean = sum(gaps) / max(1, len(gaps))
cv = (math.sqrt(sum((g - mean) ** 2 for g in gaps) / max(1, len(gaps))) / mean
if mean else 0.0)
return dict(skeleton=top_skeleton_share, out_in=out_in,
effort=high_effort, trace=trace_seek, cadence_cv=cv,
skeleton_entropy=entropy(skeletons))
def risk(f: dict) -> float:
# Placeholder weights for illustration only.
z = (3.0 * f["skeleton"] + 1.5 * min(f["out_in"] / 10, 1.0)
+ 1.0 * f["effort"] + 4.0 * f["trace"]
+ 1.0 * (1.0 if f["cadence_cv"] < 0.3 else 0.0) - 4.0)
return 1 / (1 + math.exp(-z))
Three honest caveats on this sketch. The keyword list for trace-seeking will be evaded by paraphrase, so a real system uses an embedding classifier. A low coefficient of variation in request cadence is also typical of well-behaved batch pipelines, which is why it is only one input. And the score is meant to be computed over a cluster as well, by aggregating the features of linked accounts before scoring.
Worked example with illustrative numbers
Suppose a cluster of 400 linked accounts each sends 40 requests an hour for six hours, so the cluster issues 96,000 requests. Each account alone is a modest 240 requests, comfortably below a typical limit. Now compute cluster-level features: 95 percent of prompts share one skeleton, the output-to-input ratio is 12, 90 percent of requests use the highest reasoning effort, and the cadence is nearly constant. Scored per account, a third of them might land in the medium band. Scored per cluster, every feature is extreme and the result is unambiguous. The point of the example, whose figures are invented for illustration, is that aggregation lifts a faint per-account signal above the noise floor.
Trade-offs, Gotchas, and What Goes Wrong
Hiding traces costs real users. Developers use reasoning traces to debug prompts, audit decisions and build trust in regulated settings. Summaries and hidden reasoning reduce transparency, and a number of customers will object. There is also a research cost: independent safety researchers depend on visible reasoning to study faithfulness and deception. The tension between monitorability and anti-distillation is real, and no policy resolves it fully.
False positives hurt your best customers. The users who look most like attackers, with high volume, templated prompts and heavy output, include evaluation vendors, synthetic-data teams working under a legitimate licence, and enterprise batch pipelines. Banning them by mistake is expensive. Provide an allow-listed, contractually governed path for them, and prefer throttling plus human review over automatic termination in the middle band.
Determined adversaries adapt. Once you publish a detector’s behaviour, operators will randomize cadence, vary templates with a paraphrasing model, and spread across more infrastructure. Detection is an arms race with a cost curve, not a solved problem. The aim is to make extraction slower and noisier until it stops being economical.
Summaries still leak. Even if you return only a summary, an attacker can train on millions of them, and the student will learn a coarser version of the reasoning style. Watermarks and canaries help with attribution, not prevention.
Single-provider thinking fails. Because actors probe multiple labs, a defense that only looks at your own logs misses cross-provider patterns. Participating in threat-intelligence sharing is an architectural decision, not an afterthought.
Encrypted does not mean safe. The reported incident is a clean lesson. Encryption protects confidentiality in transit and at rest; it says nothing about whether the model will decode and act on a blob presented in the wrong context. Bind secrets to the context in which they are valid, and test the replay path explicitly.
Do not trust attribution blindly. Accounts that look linked can belong to resellers, shared-VPN users or an innocent downstream customer of a bad actor. OpenAI’s own reported caveat, that it was unsure all operators were linked to one entity, is the right posture: act on behaviour, and be cautious about naming parties.
Practical Recommendations
If you operate an LLM API, start by deciding your trace exposure policy explicitly rather than inheriting a default. For most commercial reasoning tiers, that means returning an answer plus, at most, a short summary, and binding any opaque reasoning item to its originating session, user and model through authenticated encryption. Then test the replay path by deliberately feeding a blob from one session into another and confirming it is rejected.
Next, build the data foundation for detection before you build the classifier. Log per-request metadata (token counts, effort setting, timing, key and organization IDs, payment and network identifiers) in a store you can cluster on. Without that, you cannot reconstruct a campaign after the fact.
Finally, treat enforcement as a graded process with human oversight, and plan the legitimate-heavy-user path up front. Review the wider security of the model supply chain too; our piece on AI model supply chain security and provenance covers the downstream side of this problem, namely how you establish where a model came from. And if your agents call tools on untrusted input, agentic AI security and prompt injection is the companion threat model, since the same prompt-level tricks are used for extraction.
A short checklist:
- Set an explicit trace policy: raw, summary, encrypted or hidden, per model tier.
- Bind encrypted reasoning items to session, user, workspace and model.
- Scan streamed output and tool-call arguments for trace-like content.
- Gate top-tier models behind verified organizations; apply progressive limits to new accounts.
- Log features needed for rate, entropy and graph analysis; cluster by payment and infrastructure.
- Score at the cluster level and enforce in graded steps.
- Run canary prompts and consider output watermarking for attribution.
- Create a contractual path for legitimate high-volume customers.
- Join a threat-intelligence sharing channel with peers.
Frequently Asked Questions
What is a model distillation attack?
A model distillation attack is when someone uses a target model’s API outputs, at scale and against the provider’s terms, to train a competing model. Instead of stealing weights, the attacker harvests answers, and ideally hidden reasoning, then fine-tunes a cheaper student on them. It requires only API access, which is why detection rests on spotting unusual query patterns rather than on breaching any defended system.
Why do attackers want chain-of-thought traces specifically?
Chain-of-thought traces show how a model reaches an answer, including decomposition, backtracking and verification. Training a student on those steps transfers reasoning behaviour far more efficiently than training on final answers alone. Public research on open reasoning models shows small sets of curated traces can lift benchmark performance considerably, so a stolen trace is worth much more than an ordinary completion, which is why providers hide or summarize them.
Was OpenAI’s encryption broken in the reported incident?
According to press coverage of OpenAI’s statement, no. OpenAI is reported to have said the encryption was not broken, that no stored user conversations were directly accessed, and that no database was compromised. The reported weakness was that encrypted reasoning from one conversation could be replayed into another session and a model asked to decode it. OpenAI says it closed that path and strengthened protections across users, workspaces and model families.
Is distilling another company’s model illegal?
It is usually a contract violation, not clearly a crime. Major providers’ terms prohibit using outputs to build competing models, so offenders risk account termination and civil claims. Copyright protection for machine outputs is uncertain, and criminal theories such as computer misuse statutes are legally untested for API abuse. This is general information, not legal advice; the law differs by jurisdiction and is still evolving, so consult counsel for a specific situation.
How can I detect distillation on my own API?
Look for coordinated patterns, not single bad requests. Useful signals include steady round-the-clock throughput, output-heavy token ratios, high shares of maximum reasoning effort, templated prompts with near-duplicate skeletons, trace-seeking phrasing, and linked accounts sharing payment or network infrastructure. Score clusters rather than individual accounts, enforce in graduated steps, and keep a safe path for legitimate heavy users such as evaluation vendors.
Can watermarking stop distillation?
No, not alone. Token-level watermarks and canary prompts help prove after the fact that a competitor’s model learned from your outputs, which supports enforcement and disputes. But watermarks degrade under paraphrasing and mixed training data, and they do nothing to prevent extraction as it happens. Use them alongside trace controls, identity friction and behavioural detection, as one layer in a defense-in-depth design.
Further Reading
- Agentic AI security and prompt injection in 2026 for the prompt-level attack surface that extraction attempts share.
- LLM prompt caching architecture and economics for the cache side channels worth auditing.
- AI model supply chain security and provenance for establishing where a model and its training data came from.
- AI-native PLM and LLM engineering data for protecting proprietary engineering knowledge exposed to LLM workflows.
- Hinton, Vinyals and Dean, “Distilling the Knowledge in a Neural Network”.
- Kirchenbauer et al., “A Watermark for Large Language Models”.
By Riju — about
