Constrained Decoding for LLMs: Grammar-Guided Generation with XGrammar, Outlines and llguidance

Constrained Decoding for LLMs: Grammar-Guided Generation with XGrammar, Outlines and llguidance

Constrained Decoding for LLMs: Grammar-Guided Generation with XGrammar, Outlines and llguidance

Ask a language model for JSON and it will usually give you JSON. Usually is the problem. At a million calls a day, a one-in-a-hundred malformed response is ten thousand parse failures, each one a retry, a timeout or a silent data-quality bug downstream. Prompting harder, adding few-shot examples and wrapping the call in a retry loop all reduce the failure rate without ever removing it, because nothing in a sampler stops the model from emitting a stray comma.

Constrained decoding LLM techniques remove the problem at the source. Instead of asking the model to behave and checking afterwards, the inference engine edits the next-token distribution at every step so that only tokens consistent with a schema, regex or grammar can be sampled. Structured output stops being a probabilistic property of the model and becomes a structural property of the decoder.

That sounds simple, and the core idea is. The engineering is not: vocabularies hold 100,000 or more tokens, grammars are stateful, and tokens do not line up with grammar symbols. Three open-source engines, XGrammar, Outlines and llguidance, solve this in different ways and now sit underneath most serving stacks.

What this covers: the logits-masking mechanism, regex-to-finite-state-machine compilation, pushdown automata for context-free grammars, the token-boundary alignment problem, how XGrammar, Outlines and llguidance differ, integration in vLLM, SGLang and others, runtime overhead, quality side effects, and a decision matrix for choosing between engines and provider-native structured outputs.

Context and Background

Structured output has three generations. The first was prompt-level: describe the schema in words, parse the reply, retry on failure. The second was repair-level: libraries that validate against a schema and feed the error back to the model. Both treat the model as an unreliable black box and the parser as the judge. The third generation, constrained decoding, moves the judge inside the sampling loop.

The idea is older than the current LLM wave. Grammar-constrained generation has been used for semantic parsing and code synthesis for years, and the technique became mainstream for general LLM serving after Willard and Louf published “Efficient Guided Generation for Large Language Models” (arXiv 2307.09702), which recast generation as transitions through a finite-state machine and built an index over the vocabulary so that valid next tokens could be looked up rather than recomputed. That work became the Outlines library. A year later the research community split the problem into two questions: how to make masks cheap (XGrammar, llguidance) and how to make constraints non-destructive to model quality (DOMINO and the alignment literature).

Today the landscape has settled into a few layers. At the bottom are grammar engines: XGrammar from the MLC and CMU group, llguidance from the Microsoft-originated guidance project, and the Rust core behind Outlines. In the middle are serving frameworks, vLLM, SGLang, TensorRT-LLM and llama.cpp, that call those engines inside the decode loop. At the top are provider APIs such as OpenAI Structured Outputs, which by the llguidance project’s own account is powered by llguidance for JSON Schema.

For a broader architectural overview of why this matters in production pipelines, see our earlier piece on constrained decoding and structured LLM output architecture. The practical side, how often “JSON mode” actually yields valid schema-conformant output, is covered in the JSON mode and structured output benchmark. This post goes one level deeper into the machinery: what the automaton is, how the mask is computed, and where the time goes.

A note on terminology. “Guided decoding,” “structured generation,” “grammar-guided generation,” “structured outputs” and “constrained decoding” are used almost interchangeably. In vLLM the feature has been renamed more than once; the current documentation uses structured_outputs, and the older guided_json, guided_regex, guided_choice and guided_grammar request fields are documented as deprecated and removed in v0.12.0. Search older tutorials with that in mind, because their code may no longer run.

How Constrained Decoding Works: Logits Masking and Automata

Constrained decoding works by intersecting the model’s next-token distribution with the set of tokens a formal language allows. At each step the engine tracks a parser state, computes which vocabulary tokens keep the output a valid prefix, sets every other logit to negative infinity, and samples from what remains. The model never sees the constraint; only the sampler does.

Constrained decoding LLM loop: model logits are masked by a grammar automaton before sampling

Figure 1: The constrained decoding loop. The grammar engine runs beside the forward pass, produces a token bitmask for the current parser state, and the mask is applied to the logits before sampling.

Figure 1 shows the loop. The model produces a logit vector over the vocabulary. In parallel, a grammar engine holds the parser state after everything generated so far. It returns a bitmask with one bit per vocabulary token: one means “this token keeps the output a valid prefix of the language”, zero means it does not. The sampler adds negative infinity to the zero-bit logits, applies softmax, and samples. The chosen token advances the parser, and the loop repeats until the grammar reaches an accepting state and the model emits end-of-sequence.

The mask is just arithmetic on logits

Softmax turns logits z into probabilities p_i = exp(z_i) / sum over j of exp(z_j). Setting z_i to negative infinity makes exp(z_i) equal zero, so a masked token has exactly zero probability and the remaining probabilities renormalise over the allowed set. Nothing about the model changes. This also explains the memory footprint of the mask itself: a bitmask for a 128,000-token vocabulary is 128,000 bits, which is 16,000 bytes, about 16 KB per sequence per step. That is small enough to ship from CPU to GPU every step, and it is why most engines expose the mask as a packed integer tensor rather than a boolean array.

Two sampling-side subtleties matter. First, the mask must be applied before temperature scaling and top-k or top-p truncation, otherwise truncation may keep only invalid tokens and leave nothing to sample. Second, if the allowed set contains a single token, the engine can skip sampling entirely. Several engines exploit this: llguidance, for example, describes using grammars to drive “fast-forward” tokens, where deterministic stretches such as a literal key name or closing brace are emitted without waiting on a full model forward pass per token.

Where the constraint lives: regex, grammar, schema

Constraints come in a hierarchy of expressive power, and the hierarchy determines the machinery.

A choice constraint is a finite list of strings. A regular expression describes a regular language, recognised by a finite-state machine (FSM). A JSON Schema describes a nested structure that is context-free in general, because arbitrary nesting of objects and arrays needs a stack, although a schema with bounded depth can be unrolled into a regular language. A context-free grammar (CFG), commonly written in EBNF or a Lark or GBNF dialect, handles unbounded recursion such as balanced brackets, nested expressions and full programming-language syntax. Recognising a CFG needs a pushdown automaton, which is an FSM plus a stack.

Most production traffic is JSON Schema, which is why engines optimise that path hardest. Ordinary schema features compile cleanly: types, required, enum, const, items, properties. Harder features such as pattern over unbounded strings, oneOf with overlapping branches, and recursive $ref push you toward the pushdown path or are simply unsupported by a given engine. Always read the engine’s supported-keywords list, because silently ignored keywords are a common source of “valid but wrong” outputs.

Regex to FSM: the Outlines construction

The Outlines approach is easiest to understand as two compilation steps. First, a regular expression (or a JSON Schema converted into one) is compiled into a deterministic finite automaton whose alphabet is characters. Second, the engine walks the vocabulary: for every state in the automaton and for every token string in the vocabulary, it simulates feeding that token’s characters through the automaton. If the token can be consumed without hitting a dead state, the pair (state, token) is recorded along with the resulting state.

The result is an index mapping each automaton state to the allowed tokens and their destination states. At decode time there is no parsing: the engine looks up the current state, reads off the allowed token set, and after sampling jumps to the stored next state. Per-token work is a dictionary lookup. The cost has moved to compile time, which is proportional to states multiplied by vocabulary size, and that is the reason a large, complicated regex can take noticeable time to index and why the compiled index is cached.

Willard and Louf’s paper frames this as building an index over the model’s vocabulary so that guided generation adds little overhead to the token sequence generation process, and states that the approach extends to context-free grammars. The paper abstract gives no benchmark figures in the text we verified, so treat any specific speedup claim you read secondhand with caution.

Pushdown automata: why CFGs are harder

A regular language needs only a current state. A context-free language needs a stack, and that breaks the lookup-table trick. Consider JSON: after the token } the set of valid next tokens depends on what is on the stack, whether you are inside an array, inside an object after a value, or at the document root. The mask is a function of the whole stack, not one integer state. A table indexed by (state, stack) is unbounded, so it cannot be precomputed in full.

Engines handle this in three ways. Some precompute masks only for the part of the problem that does not depend on the stack, then fall back to runtime checks for the rest. Some run an Earley or GLR-style parser incrementally and compute masks on demand with a prefix tree of tokens. Some approximate by bounding stack depth. The two modern engines we cover next take the first and second approaches respectively.

XGrammar, Outlines and llguidance: Three Designs Compared

XGrammar, Outlines (via outlines-core) and llguidance all deliver the same interface, a per-step token mask, but they reach it with different trade-offs between compile-time work and runtime work. The quickest way to place them is by where the cost lands: Outlines front-loads, XGrammar splits the problem by token type, and llguidance computes lazily with a fast parser.

XGrammar token classification: context-independent tokens from a precomputed mask cache and context-dependent tokens checked at runtime with a persistent stack

Figure 2: How XGrammar splits the vocabulary. Context-independent tokens are resolved from a precomputed cache; only context-dependent tokens are checked against the live stack. The merged bitmask is applied while the GPU runs the forward pass.

XGrammar: precheck what you can, interpret what you must

XGrammar (Dong, Ruan, Cai, Lai, Xu, Zhao and Chen; arXiv 2411.15100) targets CFG execution specifically. The paper’s diagnosis is that executing a grammar normally means checking many stack states across the whole vocabulary at runtime. Its remedy has several parts, all described in the paper’s abstract and design summary:

  • Context-independent tokens are tokens whose validity can be determined from the grammar position alone, regardless of what is deeper in the stack. These are “prechecked” ahead of time and stored, so they need no runtime interpretation.
  • Context-dependent tokens are the remainder, whose validity depends on the stack contents. These must be interpreted at runtime.
  • Grammar transformations expand the grammar’s context so that more tokens become context-independent, shrinking the runtime set.
  • A persistent stack data structure speeds up the checks that remain for context-dependent tokens. Persistence here means that stack branches share structure, so exploring many candidate stack states and backtracking does not copy whole stacks.
  • Co-design with the inference engine lets grammar computation overlap with GPU execution, so mask construction hides behind the forward pass.

The abstract claims up to 100x speedup over existing solutions and describes the result as near-zero-overhead structured generation in end-to-end serving. We have not independently reproduced those figures, and the abstract gives no baseline details, so read “up to 100x” as a best-case figure against specific prior implementations rather than a typical gain.

The practical consequence of the split is that most steps are cheap. In a JSON grammar, the majority of vocabulary tokens are decided by the local grammar position: a token of pure letters is either legal inside a string or not, regardless of nesting depth. The few tokens whose legality hinges on the stack, for example a token that contains a closing bracket followed by a comma, need the live check. That is the intuition behind the design; the exact fraction depends on the grammar and tokenizer, so measure on your own schema.

According to the project README, XGrammar is the default structured generation backend for vLLM, SGLang, TensorRT-LLM and MLC-LLM, supports JSON, regex and custom context-free grammars, and runs across Linux, macOS and Windows on CPU, NVIDIA and AMD GPUs, Apple Silicon and TPUs. Python, C++, JavaScript and Swift APIs are listed. The README does not publish benchmark numbers of its own.

Outlines: index the vocabulary once

Outlines pairs the FSM index described earlier with a developer-friendly front end. In the current documentation the calling pattern is model(prompt, output_type), where the output type can be a Pydantic class, a Literal, a regex type or a JSON Schema, and models are wrapped by adapters such as outlines.from_openai. The docs’ API reference lists backend modules for outlines_core, llguidance and xgrammar, which tells you something important about the ecosystem: Outlines has become a front end that can delegate mask computation to other engines rather than a single monolithic implementation.

The strength of the index design is predictable decode-time cost for regular constraints. The weakness is the compile step: the index grows with automaton states times vocabulary, and a schema containing large enums, long literal strings or wide numeric ranges can inflate the automaton. Teams that serve many distinct schemas, as in multi-tenant APIs, must budget for compile latency on cache misses and for memory to hold the cached indices.

llguidance: Earley parsing with a token trie

llguidance, the engine behind the guidance library, is an MIT-licensed library written in Rust, with Python bindings and a C/C++ API. According to its repository, it implements a CFG parser using Earley’s algorithm on top of a lexer built from derivatives of regular expressions, and computes token masks by traversing a prefix tree of all tokens. Earley parsing handles any context-free grammar without requiring the grammar to be rewritten into a restricted form, and the derivative-based lexer avoids building large DFAs up front.

The project reports figures worth quoting with attribution: about 50 microseconds of single-core CPU time per token for a 128,000-token tokenizer with negligible startup cost; a full mask for a typical JSON schema taking about 1.5 ms before an optimisation it calls the slicer usually applies; an average under 50 microseconds across its JSON Schema Bench of 10,000 schemas and 2.5 million tokens; fewer than 1 percent of masks taking longer than 1 ms; and 0.001 percent taking longer than 10 ms. With 16 cores and a 10 ms forward pass, the project says it can handle batch sizes up to 3,200 without slowing the model. These are the maintainers’ own measurements, not independent benchmarks.

Integration dates from the repository: llama.cpp (merged 2025-02-01), SGLang (merged 2025-02-26, v0.4.4, enabled with --grammar-backend llguidance) and vLLM (merged 2025-03-25, v0.8.2). The README also states that llguidance powers OpenAI Structured Outputs for JSON Schema, shipped 2025-05-20.

Three engines compared: Outlines front-loads an index, XGrammar splits tokens by context dependence, llguidance parses lazily with Earley

Figure 3: Where each engine spends its effort. Outlines pays at compile time for a vocabulary index, XGrammar pays at compile time for a token-mask cache and at runtime for a small dependent set, llguidance pays small amounts per step through lazy Earley parsing.

Deeper Analysis: Token Alignment, Serving Integration and Overhead

Everything above assumed that grammar symbols and vocabulary tokens are the same size. They are not, and the mismatch is the subtlest source of both bugs and quality loss in constrained decoding. This section covers that alignment problem, how to call constrained decoding from vLLM with runnable code, and where the runtime overhead actually goes.

The tokenizer-versus-grammar alignment problem

A grammar speaks in characters or lexical terminals: {, "name", :, a digit. A tokenizer built with byte-pair encoding speaks in subwords: it may have one token for {", another for ":, another for "," and another for a run of spaces. The model was trained to prefer particular tokenisations of text, usually the longest-match canonical one. A mask that only permits character-by-character paths can force the model down a tokenisation it has rarely or never seen.

Beurer-Kellner, Fischer and Vechev address this in “Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation” (arXiv 2403.06988). Their diagnosis is that many existing methods add generation overhead and can reduce task accuracy because the LLM’s subword tokens do not line up with the constraint’s units. Their DOMINO algorithm enforces constraints at the subword level so token boundaries match the constraint, using pre-computation and speculative decoding to keep overhead low. The abstract reports virtually no overhead, and in some cases decoding almost twice as fast as unconstrained decoding, though it gives no figures we could verify for a specific model or schema.

Token boundary misalignment: grammar terminals versus BPE subword tokens produce distinct valid paths for the same string

Figure 4: One output string, several token paths. The grammar accepts the string however it is segmented, but only the canonical segmentation resembles what the model saw in training. Aligned engines mask at the subword level and prefer valid canonical paths.

Figure 4 shows the same JSON fragment segmented two ways. A character-level grammar accepts both. If the engine forces the character path, the model must assign probability to unusual single-character tokens, and its distribution at that point reflects training on different segmentations. The output remains syntactically valid but may be semantically worse. This is why the best engines mask over whole tokens, checking that the entire token string is a valid continuation, rather than stepping character by character.

A related trap is token healing and the prompt boundary. If your prompt ends mid-structure, for instance with an opening quote that the tokenizer would have merged with the following characters, the model sees an artificial token split it would not see in normal text. Prompt templates that end exactly at a natural token boundary avoid this. Engines differ on whether they heal the boundary automatically, so check.

Finally, special tokens and end-of-sequence handling need care. The mask must allow the end-of-sequence token only when the grammar is in an accepting state, otherwise the model can stop in the middle of an object. Conversely, if the grammar never reaches an accepting state because of a bug in the schema, generation runs until the length limit. Always set a maximum token count, and log requests that hit it.

Calling it from vLLM

vLLM is the most common place teams meet constrained decoding. Its documentation states that structured outputs generate with xgrammar or guidance backends, are enabled by default on the OpenAI-compatible server, and that the backend is selected with --structured-outputs-config.backend, whose default value auto picks a backend based on the request. Supported online constraint types are choice, regex, json, grammar and structural_tag. The offline API wraps constraints in StructuredOutputsParams and passes them through SamplingParams.

The following offline example mirrors the pattern in the vLLM documentation, extended to a JSON schema built from a Pydantic model. The choice form is taken directly from the docs; the json argument follows the documented offline constraint type, but confirm the exact parameter signature against the vLLM version you run, since this API has been renamed before.

from pydantic import BaseModel
from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams


class Ticket(BaseModel):
    customer: str
    urgency: str  # constrain further with an enum in real schemas
    summary: str


llm = LLM(model="HuggingFaceTB/SmolLM2-1.7B-Instruct")

# 1) Closed set of labels, as in the vLLM docs
labels = StructuredOutputsParams(choice=["Positive", "Negative"])
out = llm.generate(
    prompts="Classify this sentiment: vLLM is wonderful!",
    sampling_params=SamplingParams(structured_outputs=labels, max_tokens=8),
)
print(out[0].outputs[0].text)

# 2) JSON constrained by a schema
schema = Ticket.model_json_schema()
js = StructuredOutputsParams(json=schema)
out = llm.generate(
    prompts="Extract a ticket: Alice cannot log in and needs help ASAP.",
    sampling_params=SamplingParams(structured_outputs=js, max_tokens=200),
)
ticket = Ticket.model_validate_json(out[0].outputs[0].text)
print(ticket)

Two operational notes. The max_tokens limit is a safety net, not a style choice: a truncated JSON object is invalid and model_validate_json will raise. And validate even though the output is guaranteed structurally valid, because the guarantee covers the grammar you supplied, not business rules such as “urgency must match the ticket text”.

On the online server the same constraint goes through extra_body={"structured_outputs": {...}} or through response_format with a JSON schema. The deprecated guided_* fields and guided_decoding_backend were documented as removed in v0.12.0, so migrate any older client code.

Other serving stacks

SGLang accepts a grammar backend flag, with llguidance enabled by --grammar-backend llguidance per the llguidance repository, and lists XGrammar among its backends per the XGrammar README. TensorRT-LLM and MLC-LLM also use XGrammar by default according to the same README. llama.cpp supports GBNF grammars natively and can use llguidance when built with the -DLLAMA_LLGUIDANCE=ON option. We did not verify per-version flag spellings for TensorRT-LLM or MLC-LLM this run, so check their current release notes before copying configuration.

Where the per-token overhead goes

Overhead in constrained decoding has four components, and conflating them produces bad capacity plans.

  1. Compile time. Turning a schema or grammar into an automaton or mask cache. It is paid once per distinct constraint, so it dominates when every request carries a unique schema and vanishes when schemas repeat. Cache compiled grammars keyed by schema hash.
  2. Mask computation. The per-step CPU work to produce the bitmask. The llguidance numbers above, tens of microseconds on average with a long tail, show that the average is not the number to plan around. The tail matters because one slow mask stalls the whole batch step.
  3. Mask transfer and application. Moving a roughly 16 KB bitmask per sequence per step to the GPU and applying it to logits. For a batch of 256 sequences that is about 4 MB per step, small relative to weights but not zero, and it adds a synchronisation point.
  4. Lost parallelism. Speculative decoding, multi-token prediction and batching all interact with per-sequence grammar state. A speculated draft token must be validated against the grammar, and a rejected draft requires rolling the parser back. Engines that support rollback, such as through a persistent stack, make this cheap.

A worked, deliberately illustrative budget helps. Suppose a decode step takes 10 ms of GPU time, which is the figure the llguidance README uses when it reasons about batch capacity. If mask computation for one sequence averages 50 microseconds on one core, then 200 sequences need 10 ms of single-core CPU time, which equals the whole step on one core, so a handful of cores or an overlapped design is needed to stay hidden behind the forward pass. That arithmetic is the reason XGrammar’s co-design claim, overlapping grammar work with GPU execution, and llguidance’s multi-core claim matter more than the headline per-token figure.

The honest answer to “what is the overhead” is therefore: for common JSON schemas on modern engines it is small enough to hide behind the forward pass, at the mean; the tails, first-request compile latency, and exotic grammars are where it hurts. Measure time-to-first-token and inter-token latency with and without the constraint on your own schemas. Our vLLM, TGI, SGLang and Triton inference benchmark describes a methodology you can adapt for that comparison.

Quality, Distribution Distortion and Provider-Native Structured Outputs

Constrained decoding guarantees syntax, not truth. Whether it also preserves quality is a separate empirical question, and the answer is “usually, but not always”.

What masking does to the distribution

Masking and renormalising samples from the model’s distribution conditioned on the output being valid, one token at a time. That is greedy local conditioning. It is not the same as sampling from the model’s distribution over complete valid sequences. The difference is called the myopia or distribution-distortion problem: a token that is locally valid may lead into a region where all continuations are low-probability for the model, and the renormalisation at each step hides that. In the worst case the constraint steers the model into a valid but nonsensical object, because every locally cheap valid choice was taken.

In practice distortion is small when the model’s own preference already agrees with the constraint, as with a well-instructed model asked for JSON, and large when the constraint contradicts what the model wants to say, for example forcing a one-word label from a closed set when the model would have explained itself.

Reasoning degradation when over-constrained

Tam and colleagues, in “Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models” (arXiv 2408.02442), report that requiring structured output such as JSON or XML significantly reduces reasoning performance compared with free-form responses, and that stricter format constraints generally cause larger drops. We could verify the headline finding from the abstract but not the specific models, tasks or magnitudes. Whether the effect holds across models and tasks is something to test on your own workload, so treat it as a caution about schema design rather than a law.

The mechanism is plausible from first principles. If the schema begins with an answer field and the model must commit to it before any reasoning tokens, the model loses its chain of thought. The fix is structural, not mathematical: put a reasoning string field before the answer field, or let the model think in free text first and constrain only the final extraction. Field order in the schema is order of generation, and that is a design decision, not an accident.

A second lever is constraining as little as needed. Prefer enum for genuinely closed sets, but allow free strings with a length cap where the content is open. Avoid regexes that exclude natural phrasing, and avoid forcing exact numeric precision that the model cannot know.

Constrained decoding versus native provider structured outputs

Hosted APIs now offer their own structured output modes. The llguidance repository states that OpenAI Structured Outputs, shipped on 2025-05-20, is powered by llguidance and covers JSON Schema only. The consequences for a practitioner:

  • Hosted, JSON Schema only: you get a guarantee with no infrastructure, but you accept the provider’s supported-keyword subset, schema size limits, and first-request latency for new schemas. You cannot supply an arbitrary CFG.
  • Self-hosted engine: you can use regex, full CFGs, tool-call grammars, and domain languages such as SQL dialects, and you control caching, but you operate the stack and own the compatibility testing between engine, tokenizer and model.
  • Mixed estates: many teams use a hosted API for general extraction and a self-hosted engine where the output language is custom or the data cannot leave the network.

Provider limits change quickly, so we deliberately give no numeric schema limits here. Check the provider’s current structured-output documentation for supported keywords, nesting depth and property counts before designing around them.

Decision matrix

The table summarises which tool fits which constraint. Entries describe design tendencies drawn from the sources above, not measured rankings.

Need XGrammar Outlines (outlines-core) llguidance Provider-native
Simple JSON Schema at scale Strong, default in several servers Good, index cached per schema Strong, low per-step cost Easiest, no infra
Regex or closed choice Supported Original strength Supported through lexer Usually not exposed
Arbitrary CFG or DSL Supported Supported Supported, Earley handles general CFGs Generally not available
Many unique schemas per minute Compile cost per new grammar Compile cost grows with automaton size Lazy, small startup cost Provider-side first-call latency
Deep recursion Persistent stack helps Needs pushdown support Earley handles it Subset only
Operations burden Medium, bundled in server Medium, library plus adapter Medium, bundled or linked Lowest

Trade-offs, Gotchas, and What Goes Wrong

The failure modes of constrained decoding are quieter than a parse error, which is why they survive into production.

Valid is not correct. The guarantee is that output parses against your grammar. A schema with "type": "string" for a date accepts “tomorrow-ish”. Encode as much as the grammar can express, use enums, formats and bounded lengths, and still validate semantically afterwards.

Silently dropped schema keywords. Engines support a subset of JSON Schema. If pattern, minimum, uniqueItems or oneOf is ignored rather than rejected, the model can return output that satisfies the engine’s subset but violates your intent. Run a conformance test: generate a few hundred outputs and validate them with a full validator such as the jsonschema library.

Dead ends and infinite loops. A grammar can allow whitespace indefinitely, and a model with a tendency to emit newlines will fill the budget with blanks. vLLM’s older guided_whitespace_pattern option existed for this reason. Bound whitespace in the grammar, use compact formatting, and always set max_tokens. Likewise a schema with an unreachable accepting state hangs generation until the limit.

Unbounded arrays and strings. An array with no maxItems lets the model keep adding entries, and a long list that hits the token limit is truncated mid-object and unparseable. Bound lengths and size max_tokens to the worst legitimate output.

Compile latency cliffs. The first request carrying a new large schema pays compilation. Pre-warm common schemas at deploy time, and put an upper bound on schema size at your API gateway. A user-supplied schema is also a denial-of-service vector: a pathological regex or a deeply nested grammar can consume CPU, so treat externally supplied grammars as untrusted input.

Tokenizer drift. Masks are computed against a specific tokenizer. If you swap the model or tokenizer version, cached grammars keyed only by schema hash become wrong. Include the tokenizer identity in the cache key.

Speculative decoding and batching interactions. Combining grammars with speculative decoding is possible but each engine and serving version has its own support matrix. Verify with a test, not an assumption, that quality and validity hold with your exact flags enabled.

Reasoning-before-answer. As covered earlier, schema field order determines whether the model gets to think. An answer-first schema can quietly cost accuracy on multi-step tasks.

Security. Constraining structure does not constrain content. A model forced to emit valid JSON can still emit a prompt-injected tool argument. Treat decoded values as untrusted data and validate them before acting.

Practical Recommendations

Start from the output language, not the tool. If your need is a JSON object in a hosted API call, use the provider’s native structured output and stop there. If you self-host, use the engine your server already bundles: XGrammar or guidance in current vLLM per its documentation, and whichever backend your SGLang or TensorRT-LLM version defaults to. Reach for regex or CFG constraints only when JSON Schema cannot express the format.

Design the schema for the model, not just the parser. Put reasoning fields before answer fields, give fields descriptive names, bound array lengths and string sizes, use enums for closed sets, and avoid deep nesting you do not need. Keep a plain-text escape hatch for cases where the model cannot comply, for example a nullable field or an unknown enum value, so it is not forced to fabricate.

Treat the feature as infrastructure that needs testing. Build a regression suite of schemas and prompts, run it on every engine, model and tokenizer upgrade, and track three metrics: schema-validity rate (should be 100 percent), semantic accuracy (the number that actually matters) and added latency at the median and the 99th percentile.

Checklist before shipping:

  • [ ] Engine and backend pinned, and the flag names verified against the version in production
  • [ ] Compiled grammars cached, keyed by schema hash and tokenizer
  • [ ] Common schemas pre-warmed at deploy time
  • [ ] max_tokens set from the worst legitimate output, with truncation alerts
  • [ ] Whitespace and array lengths bounded in the schema
  • [ ] Conformance test using a full JSON Schema validator
  • [ ] Semantic validation after decoding, and decoded values treated as untrusted
  • [ ] A/B accuracy comparison against unconstrained generation with a parse-and-retry baseline
  • [ ] External or user-supplied grammars size-limited and sandboxed

If your workload spans signals and decoding of a very different kind, such as translating neural activity into discrete outputs under hard constraints, the same mask-then-sample logic appears in brain-computer interface neural decoding architectures, where language-model priors constrain a noisy decoder.

Frequently Asked Questions

What is constrained decoding in LLMs?

Constrained decoding restricts which tokens a language model may emit at each step so the output always conforms to a schema, regular expression or grammar. The inference engine tracks a parser state, computes a bitmask of valid next tokens, sets invalid logits to negative infinity and samples from what remains. The model itself is unchanged; the guarantee comes from the sampler, not from training or prompting.

How is constrained decoding different from JSON mode?

JSON mode usually guarantees only that the output is syntactically valid JSON, and sometimes not even that. Constrained decoding against a JSON Schema also enforces keys, types, enums and nesting, so the object matches your structure. Regex and full grammar constraints go further, covering formats JSON cannot express. Naming varies by provider, so check whether a mode enforces a schema or only well-formed JSON.

Does constrained decoding slow down inference?

It adds mask computation, mask transfer and some loss of parallelism. Engine maintainers report small costs: the llguidance project cites about 50 microseconds per token on average for a 128,000-token tokenizer, and the XGrammar paper claims near-zero overhead with up to 100x speedup over earlier solutions. These are vendor-reported. Compile latency on new schemas and slow tail masks are the real risks, so benchmark your own schemas.

Does constrained decoding reduce answer quality?

It can. Masking conditions on validity one token at a time, which can distort the distribution, and the paper “Let Me Speak Freely?” reports that strict format restrictions can reduce reasoning performance. The usual remedy is schema design: place a free-text reasoning field before the answer, constrain as little as necessary, and compare accuracy against an unconstrained baseline on your own task.

Which is better, XGrammar, Outlines or llguidance?

None dominates; they optimise different costs. XGrammar splits tokens into context-independent and context-dependent sets and is the default backend in several servers per its README. Outlines front-loads an FSM vocabulary index and offers a convenient Python API. llguidance is a Rust Earley-based engine with low per-token cost. Pick the one your serving stack supports well, then measure on your schemas.

Can I use constrained decoding with a hosted API?

You can use the provider’s native structured output feature, which for JSON Schema is reported to be powered by llguidance in OpenAI’s case. You usually cannot supply arbitrary grammars or regexes. For custom formats such as a domain-specific language, or when data must stay on your network, self-host a model with vLLM, SGLang or llama.cpp and a grammar backend.

Further Reading

Not professional advice; verify engine flags and versions against current documentation before production use.

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *