LLM Logprobs Explained: Token Probabilities in Practice
Logprobs are the log probabilities an LLM assigns to each token it generates. What they mean, how to read them, and practical uses like classification.
Logprobs are the logarithms of the probabilities a large language model assigns to tokens as it generates text. At every step, the model produces a probability for every token in its vocabulary; the logprob of the chosen token tells you how likely the model considered it, and many APIs can also return the top few alternatives at each position. They are the closest thing an LLM offers to a built-in confidence signal, and they unlock practical techniques for classification, ranking, and quality checks.
Where logprobs come from
A language model generates one token at a time. For each position, the final layer outputs a vector of raw scores called logits, one per vocabulary entry. A softmax turns those logits into a probability distribution that sums to 1. The decoder then picks a token from that distribution according to the sampling settings, as covered in our guide to temperature, top-p, and top-k.
The logprob is simply the natural logarithm of the probability for a given token:
- A probability of 1.0 gives a logprob of 0.
- A probability of 0.5 gives about −0.69.
- A probability of 0.1 gives about −2.3.
- A probability of 0.01 gives about −4.6.
So logprobs are always zero or negative, and closer to zero means more confident. To get back to a probability, exponentiate: p = exp(logprob).
Why logs instead of probabilities?
Two reasons, both practical. First, numerical stability: probabilities of long sequences become vanishingly small, and multiplying many small numbers underflows floating-point precision. Second, convenience: the probability of a sequence is the product of its token probabilities, and the log of a product is a sum. Adding logprobs is far easier and safer than multiplying probabilities, a theme familiar to anyone who has worked with floating-point arithmetic.
What an API response looks like
When logprobs are enabled, each generated token comes back with its logprob and, optionally, a short list of the most likely alternatives at that position. Conceptually:
{
"token": " Paris",
"logprob": -0.02,
"top_logprobs": [
{ "token": " Paris", "logprob": -0.02 },
{ "token": " Lyon", "logprob": -4.1 },
{ "token": " France", "logprob": -5.3 }
]
}
Here the model was about 98% sure of ” Paris”. Note the leading space: tokens often include whitespace, which matters when you compare strings. Field names and how many alternatives you can request vary by provider, and not every model or endpoint exposes logprobs at all.
Practical uses
Classification with calibrated scores
Ask the model to answer with a single label, such as positive, negative, or neutral, and read the logprobs of each label at the first output position. Instead of a bare answer, you get a probability distribution over classes. You can then set thresholds, route low-confidence cases to a human, or plot precision against recall. This is often more useful than asking the model to “rate its confidence from 1 to 10,” which produces numbers that are poorly grounded.
Constrain the output tightly so each label maps to a single, distinct first token. Our guide to structured outputs and JSON mode covers ways to force a fixed format.
Detecting uncertainty and likely hallucinations
Low-probability tokens in factual positions, such as a date, a name, or a number, are a useful warning sign. A model that is unsure about a citation will often show it in the logprobs even when the prose sounds fluent. This is not a reliable hallucination detector on its own, since models can be confidently wrong, but it is a cheap signal for flagging outputs that deserve review. See LLM hallucinations explained for the broader picture.
Ranking and reranking
Given a query and several candidate answers or documents, you can score each candidate by the total or average logprob the model assigns to it. Average per-token logprob avoids penalising longer candidates simply for being longer. This technique underlies some reranking steps in retrieval-augmented generation pipelines.
Perplexity and evaluation
Perplexity is the exponential of the negative average logprob over a text. Lower perplexity means the model found the text more predictable. It is a standard intrinsic metric for comparing language models on the same tokenizer, and a building block in many LLM eval setups.
Autocomplete and UI hints
Interfaces can use top alternatives to offer suggestions, highlight uncertain spans for the user to check, or decide whether to show a completion at all.
Interpreting logprobs carefully
A few caveats keep logprobs from being misused:
- They reflect the model’s distribution, not ground truth. A high probability means the model expected that token, not that the statement is correct.
- Calibration varies. Base models are often reasonably calibrated; models tuned heavily with human feedback can become overconfident, concentrating probability on one answer. Check calibration on your own labelled data before relying on thresholds.
- Sampling settings matter. Depending on the provider, returned logprobs may describe the raw distribution or the distribution after temperature and truncation are applied. Know which one you are reading.
- Tokenization splits words. A word like “unbelievable” may be several tokens. Its probability is the product of all of them, so sum their logprobs, and be aware that alternative tokenizations of the same string exist.
- Only top-k alternatives are visible. If a label is not in the returned top list, you only know its logprob is lower than the smallest one shown.
Logprobs vs other confidence signals
| Signal | What it measures | Strengths | Weaknesses |
|---|---|---|---|
| Logprobs | Model’s token-level probability | Cheap, granular, no extra calls | Calibration varies; not always exposed |
| Self-reported confidence | Model’s verbal estimate | Works with any model | Often poorly calibrated |
| Self-consistency | Agreement across several samples | Model-agnostic, robust | Several times the cost |
| External verifier | Separate model or check | Can catch confident errors | More infrastructure |
The takeaway
Logprobs are the log of the probability a model assigned to each token it produced, with values closer to zero meaning higher confidence. They are cheap to request where supported and turn an LLM’s raw distribution into something you can threshold, rank, and evaluate. Use them for classification scores, uncertainty flags, and reranking, but treat them as the model’s belief rather than a guarantee of correctness, and verify calibration on your own data.
Tagged
Keep reading
Chisato · · 5 min read Constrained Decoding: How LLMs Output Guaranteed JSON
Constrained decoding masks invalid tokens at each step so an LLM can only emit output matching a grammar, regex, or JSON Schema. How it works and its limits.
Chisato · · 4 min read Precision vs Recall, Explained
Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.
Chisato · · 4 min read What Is Logit Bias? Steering LLM Output Per Token
Logit bias nudges an LLM's token probabilities up or down before sampling, letting you ban, force, or discourage specific words without a prompt.