Articles

Constrained Decoding: How LLMs Output Guaranteed JSON

Constrained decoding masks invalid tokens at each step so an LLM can only emit output matching a grammar, regex, or JSON Schema. How it works and its limits.

Chisato Chisato · · 5 min read
Glowing purple fibers resembling a neural network

Constrained decoding is a technique that guarantees a language model’s output matches a formal structure, such as a JSON Schema, a regular expression, or a context-free grammar, by blocking every token that would break that structure before the model is allowed to pick one. Instead of asking the model nicely to return valid JSON and hoping, the decoder makes invalid output impossible to generate.

It’s the mechanism underneath most “strict” structured output and JSON mode features. Understanding how it works explains both why those features are reliable and why they occasionally produce strange results.

How normal decoding works

A language model generates text one token at a time. At each step it produces a score, or logit, for every token in its vocabulary. Those scores become a probability distribution, and a sampling strategy picks the next token: greedy selection, or methods shaped by temperature, top-p, and top-k. The chosen token is appended, and the loop repeats.

Nothing in that loop knows about JSON. If the model is well trained and well prompted, it will usually produce valid output. “Usually” is the problem: a missing brace, an unescaped quote, or a stray sentence before the opening { is enough to break a parser in production.

The core trick: masking logits

Constrained decoding adds one step to the loop. Before sampling, it computes which tokens are legal given the output so far and the target structure, and sets the logits of every illegal token to negative infinity. Those tokens now have zero probability. The sampler then picks from whatever remains.

  1. Track where the output currently is in the grammar.
  2. Compute the set of tokens that can legally come next.
  3. Mask every other token.
  4. Sample from the masked distribution.
  5. Advance the grammar state with the chosen token and repeat.

This is the same lever as logit bias, applied automatically and exhaustively at every step. If the grammar says the next thing must be a closing quote or a digit, those are the only tokens the model can choose.

Turning a schema into a state machine

Checking every vocabulary token against a grammar at every step would be far too slow, since vocabularies contain tens of thousands of tokens or more. Practical implementations compile the constraint ahead of time.

For constraints expressible as regular expressions, which covers a surprising amount of JSON Schema (fixed keys, enums, string patterns, numbers, bounded nesting), the constraint becomes a finite state machine. The same theory behind how regex engines work applies: a deterministic automaton where each state represents “how much of the pattern has been matched so far.”

The key optimization is precomputing, for each automaton state, which vocabulary tokens are allowed. At generation time, finding the mask becomes a lookup rather than a search. The schema is compiled once and reused across every request that uses it.

Structures with arbitrary nesting, such as recursive JSON objects, can’t be captured by a finite automaton. These need a pushdown automaton, essentially a state machine with a stack, so the decoder can track how many brackets are open. That’s more expensive to evaluate per step, and implementations use caching and partial precomputation to keep it fast.

The tokenization wrinkle

Grammars are defined over characters. Models generate tokens, and a single token can span several characters, including characters that cross grammar boundaries. A token like ":" or "},{ might complete a string, emit a separator, and open a new object in one move. Understanding how tokenization works makes it clear why this is tricky.

The decoder therefore has to ask, for each token, “if I walk the entire character sequence of this token through the automaton from the current state, do I end in a valid state?” Precomputed indexes handle this, but it’s also the source of subtle quality issues. The model might prefer to write a value using a token boundary that the constraint forbids, forcing it down a less natural tokenization it rarely saw during training.

Prompting vs constrained decoding

Prompt-only structured outputConstrained decoding
Validity guaranteeBest effortGuaranteed for the specified structure
Needs output validationYes, alwaysOnly for semantic checks
Setup costNoneSchema compilation, often cached
Supported structuresAnything you can describeLimited to what the engine can compile
Effect on content qualityModel writes naturallyCan distort output if constraints are awkward

What constrained decoding doesn’t guarantee

Constrained decoding guarantees syntax, not truth. The output will parse and match your JSON Schema, but every field can still be wrong. A model forced into an enum will always pick one of its values, even when none fits. A required field will always be filled, even if the model has to make something up.

There are a few practical failure modes to know about:

  • Forced commitment. If the schema requires an answer field before any reasoning field, the model must commit to an answer before it has “thought” in text. Ordering fields so explanation comes before the conclusion often improves results.
  • Truncation. If the token limit is reached mid-object, you get structurally incomplete output. The constraint can’t close brackets it never got to emit.
  • Endless strings. An unbounded string or array field gives the model room to ramble. Length limits in the schema help.
  • Unsupported keywords. Many engines support only a subset of JSON Schema. Complex conditionals or cross-field rules may be ignored or rejected rather than enforced.
  • Distribution shift. Heavy constraints can push the model toward tokens it considers unlikely, which can lower quality compared with letting it write freely and validating after.

Where it fits in an application

Constrained decoding is most valuable where a downstream program consumes the output directly: function calling, extraction pipelines, classification into fixed labels, and agent tool arguments. It eliminates an entire class of retry loops caused by malformed output.

It doesn’t replace validation of meaning. Check that IDs exist, numbers fall in range, and dates make sense, the same way you’d validate any untrusted input. Treat the structure as guaranteed and the content as a claim.

The takeaway

Constrained decoding enforces structure by masking every token that would violate a grammar, so the model literally cannot produce invalid output. Schemas are compiled into state machines with precomputed token masks, which keeps it fast enough for production. It guarantees well-formed output, not correct output, and awkward schemas can hurt quality. Design schemas that let the model reason before it answers, keep fields bounded, and keep validating the content.

Chisato Chisato · · 4 min read

Precision vs Recall, Explained

Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.

#AI #Machine Learning #LLMs
Chisato Chisato · · 4 min read

What Is Logit Bias? Steering LLM Output Per Token

Logit bias nudges an LLM's token probabilities up or down before sampling, letting you ban, force, or discourage specific words without a prompt.

#AI #LLMs #Machine Learning