What Is Test-Time Compute? Inference-Time Scaling
Test-time compute is extra computation an AI model spends while answering, not while training — trading latency and cost for better answers.
Test-time compute is the computation a model spends after training, while it’s actually generating an answer — as opposed to the compute spent building the model in the first place. For years, the main lever for improving a model was training-time scale: more parameters, more data, more GPU-hours before the model ever answered a question. Test-time compute is a second, independent lever: keep the trained model fixed, but let it “think” longer, explore more paths, or check its own work before returning a response.
Why a second scaling axis matters
Training-time scaling has a straightforward story: bigger models trained on more data tend to perform better, following predictable scaling laws. But training is a one-time cost amortized over every future query, while inference happens on every single request. If a task is easy, spending a fixed, large amount of compute on it is wasteful. If a task is hard — a multi-step proof, a tricky debugging session, a long planning problem — a small model given more time to reason can sometimes match or beat a much larger model that answers instantly.
This is the core trade test-time compute makes: latency and inference cost in exchange for accuracy, on a per-query basis rather than a fixed-at-training-time basis. It’s the same trade a person makes when they stop to double-check arithmetic instead of blurting out the first number that comes to mind.
How models spend extra time at inference
There are a few common mechanisms, often combined:
- Longer chain-of-thought. The model generates intermediate reasoning steps before its final answer, effectively “thinking out loud.” See chain-of-thought prompting explained for how this works at the prompting level; reasoning models bake extended reasoning into the model itself rather than relying on prompt tricks.
- Sampling multiple candidates. Generate several independent answers and pick the best one, either by majority vote or with a separate scoring model. This is a form of parallel test-time compute rather than sequential.
- Search over the reasoning tree. Instead of committing to one line of reasoning, explore several branches and prune the weaker ones — conceptually related to beam search, extended from token-level search to reasoning-step-level search.
- Self-verification. The model (or a second model acting as a critic) checks intermediate steps or the final answer for consistency before returning it, and revises if something doesn’t add up.
Each of these consumes more tokens, and tokens are what you pay for and wait for. A model using heavy test-time compute might generate thousands of “thinking” tokens the user never sees before producing a short final answer.
The trade-offs
More test-time compute is not free, and it’s not always the right choice.
| Low test-time compute | High test-time compute | |
|---|---|---|
| Latency | Fast, near-instant | Seconds to minutes |
| Cost per query | Low | Can be substantially higher |
| Accuracy on easy tasks | Sufficient | Marginal gain, wasted spend |
| Accuracy on hard, multi-step tasks | Often insufficient | Meaningfully better |
| Predictability | Consistent response time | Variable, harder to budget for |
Because cost scales with tokens generated, teams building on top of reasoning models often route queries: cheap, low-effort model calls for straightforward requests, and a heavier reasoning pass reserved for genuinely hard ones. Our LLM token cost calculator is useful here — it makes the cost difference between a short direct answer and a long reasoning trace concrete, which helps when deciding whether a given feature needs a reasoning model at all.
How it differs from other efficiency levers
It’s worth distinguishing test-time compute from a few adjacent concepts that also affect inference cost and speed:
- Quantization reduces the numerical precision of a model’s weights to make each forward pass cheaper — see what is quantization. It shrinks the cost per step; test-time compute increases the number of steps.
- Speculative decoding speeds up generation of a fixed amount of output by using a small draft model to propose tokens a larger model then verifies — see what is speculative decoding. It’s about generating the same output faster, not generating more reasoning.
- The KV cache makes each additional token in a long context cheaper to generate by avoiding recomputation — see what is a KV cache. It supports long reasoning traces efficiently but isn’t what causes a model to reason longer in the first place.
- Mixture of experts changes how many of a model’s parameters are active per token, affecting training and inference cost together — see what is a mixture of experts. It’s a training-time architecture choice, largely orthogonal to test-time reasoning length.
Test-time compute sits alongside these as a separate knob: given a fixed, already-trained (and possibly already-quantized) model, how much extra work should it do per query before answering.
Why this shifted how labs think about scaling
The practical upshot is that “how good is this model” is no longer a single number tied only to its parameter count and training data. It’s a curve: accuracy as a function of how much compute you’re willing to spend at answer time. A smaller model allowed to reason extensively can land on the same curve as a larger model answering immediately, at a different point of the latency-cost-accuracy trade-off. That reframes model selection from “pick the biggest model you can afford” to “pick the model and reasoning budget combination that fits this specific task.”
It also changes capacity planning for anyone serving these models. A fixed-cost, low-latency chat assistant and a highly variable, occasionally slow reasoning agent working through a multi-step task behave very differently under load, even if they’re built on related models. Systems that mix both need to size infrastructure for the reasoning tail, not just the average case, and often benefit from batching decisions between batch and real-time inference depending on how latency-sensitive the workload actually is.
The takeaway
Test-time compute is the amount of computation a model spends while generating an answer, distinct from the compute spent training it. Techniques like extended chain-of-thought, sampling multiple candidates, and self-verification all trade latency and cost for accuracy on a per-query basis, which is especially valuable for hard, multi-step problems and largely wasted on easy ones. Treat it as a separate dial from model size, quantization, or decoding speed — and route queries so that expensive reasoning is spent only where it actually changes the outcome.
Tagged
Keep reading
Chisato · · 5 min read Constrained Decoding: How LLMs Output Guaranteed JSON
Constrained decoding masks invalid tokens at each step so an LLM can only emit output matching a grammar, regex, or JSON Schema. How it works and its limits.
Chisato · · 5 min read LLM Logprobs Explained: Token Probabilities in Practice
Logprobs are the log probabilities an LLM assigns to each token it generates. What they mean, how to read them, and practical uses like classification.
Chisato · · 4 min read Precision vs Recall, Explained
Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.