Test-Time Compute Scaling: Reasoning Budgets, Best-of-N and Process Reward Models

Test-Time Compute Scaling: Reasoning Budgets, Best-of-N and Process Reward Models

Test-Time Compute Scaling: Reasoning Budgets, Best-of-N and Process Reward Models

For most of the last decade, getting a better answer out of a language model meant training a bigger one. That assumption quietly broke. A model that is frozen, deployed and unchanged can now be made markedly more accurate on hard problems simply by letting it think for longer, sample more candidates, or search against a verifier. Test-time compute scaling is the name for that idea: spend more floating-point operations at inference, per question, and trade them for accuracy.

It matters now because every major provider ships a control for it. Anthropic exposes a thinking budget and an effort setting, OpenAI exposes a reasoning effort level, and Google exposes thinking levels. Your bill, your latency and your error rate are all downstream of how you set those dials. Most teams set them once and never revisit the choice.

This post gives you the mental model and the arithmetic to set them deliberately. You will leave knowing which technique fits which problem, how to compute cost per correct answer, and where the returns flatten.

What this covers: the research lineage, the two axes of scaling (parallel and sequential), verifiers and process reward models, search, provider budget APIs, worked cost arithmetic, a runnable Python sketch, failure modes and a decision matrix.

Context and Background

The classical scaling story, popularised by the Kaplan and Chinchilla papers, says loss falls predictably with parameters, data and training compute. It is a pre-training story. Inference was treated as a fixed, cheap cost: one forward pass per token, one answer per prompt.

Three lines of work cracked that framing open. The first is repeated sampling and voting. Wang et al.’s self-consistency paper (Xuezhi Wang, Jason Wei, Denny Zhou and colleagues) replaced greedy chain-of-thought decoding with sampling several diverse reasoning paths and picking the answer that most paths agree on. Its abstract reports gains over standard chain-of-thought prompting of 17.9 percent on GSM8K, 12.2 percent on AQuA, 11.0 percent on SVAMP, 6.4 percent on StrategyQA and 3.9 percent on ARC-challenge.

The second line is verifier-guided selection. Cobbe et al., in “Training Verifiers to Solve Math Word Problems”, introduced the GSM8K dataset of 8.5K grade-school problems and trained verifiers to judge candidate solutions. At test time the system generates many candidates and returns the one the verifier ranks highest. The authors report that verification beats a finetuning baseline and scales better with more data.

The third line is the process reward model. Lightman et al., “Let’s Verify Step by Step” (OpenAI), compared rewarding only the final answer (outcome supervision) with rewarding every intermediate step (process supervision). They released PRM800K, 800,000 step-level human labels, and reported that their process-supervised model solved 78 percent of problems from a representative subset of the MATH test set.

Then came the synthesis. Snell, Lee, Xu and Kumar published “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters” in August 2024 (arXiv:2408.03314). They studied two mechanisms, searching against dense process-based verifiers and adaptively updating the model’s response distribution per prompt, and found that the best choice depends on prompt difficulty. Their compute-optimal strategy improved test-time efficiency by more than 4x over a best-of-N baseline, and in a FLOPs-matched evaluation a smaller model with test-time compute could beat a model 14x larger on problems where the small model already had non-trivial success.

That last caveat is the part most summaries drop. The trade is not free and not universal. If you want the product-level view of how a flagship reasoning model packages this, see our analysis of OpenAI o3 and reasoning-model test-time compute, and for the budgeting side see AI inference cost optimization.

The Core Idea: Two Axes and a Controller

Test-time compute scaling is the practice of allocating extra inference computation per query, through longer reasoning, more sampled candidates, or search against a verifier, to raise accuracy without changing the model weights. The gain depends on problem difficulty, verifier quality and how the extra compute is spent, and it flattens as budgets grow.

Think of any scheme as three parts: a generator that proposes reasoning, a selector that decides which proposal to trust, and a controller that decides how much to spend. The controller is the part most production systems lack, and it is where Snell et al. found the largest efficiency gain.

Test-time compute scaling controller routing prompts by difficulty to single pass, parallel sampling or sequential revision

Figure 1: A compute-optimal controller routes each prompt by estimated difficulty, then meters the spend.

Figure 1 shows the shape. Easy prompts take a cheap single pass. Medium prompts get parallel samples with a vote or verifier. Hard prompts get long sequential reasoning or revision. Every path ends in a cost meter, because a scaling policy you cannot measure is a policy you cannot tune.

Parallel scaling: sample many, then select

Parallel scaling draws N independent chains from the same model at non-zero temperature. The chains do not see each other, so the work is embarrassingly parallel and latency stays close to that of a single chain if you have the throughput. You then need a selector.

The simplest selector is majority vote over final answers, which is self-consistency. It needs no extra model, only a way to extract and compare final answers, so it fits tasks with short, canonical answers such as numbers, multiple-choice letters or normalised strings. It fails on open-ended text, where no two outputs are identical.

The next selector is a learned verifier that scores each chain, and you pick the highest-scoring one. This is best-of-N. It works on open-ended tasks, but it is only as good as the verifier, and it is where reward hacking enters, which we cover below.

The ceiling for parallel scaling is the coverage of the model: the fraction of problems for which at least one sample is correct. Brown et al., “Large Language Monkeys” (arXiv:2407.21787), measured it. They report that coverage scales with the number of samples across four orders of magnitude, often log-linearly, and that with DeepSeek-Coder-V2-Instruct the share of SWE-bench Lite issues solved rose from 15.9 percent with one sample to 56 percent with 250 samples. The same paper notes that majority voting and reward models stop improving beyond several hundred samples. Coverage can keep climbing while selection cannot keep up. That gap is the real limit of parallel scaling.

Sequential scaling: think longer, revise, retry

Sequential scaling lets one chain run longer: more reasoning tokens, explicit self-checking, backtracking, or revising an earlier draft conditioned on its own mistakes. Each token depends on the previous ones, so latency grows roughly linearly with the budget. It cannot be parallelised across the chain, though a single chain still benefits from fast decoding hardware.

Reasoning models are trained to do this natively. Their post-training teaches them to emit long internal chains that explore, check and correct. The s1 paper (Muennighoff et al., arXiv:2501.19393) showed how little machinery is needed to expose the dial. It fine-tuned Qwen2.5-32B-Instruct on s1K, a curated set of 1,000 questions with reasoning traces, then applied budget forcing: cut the model off to end thinking early, or append the word “Wait” to push it to keep checking. The authors report that scaling with budget forcing raised AIME24 accuracy from 50 percent to 57 percent, beyond what the model reached without the intervention.

The practical point is that sequential scaling has an internal selector: the model’s own self-verification. That makes it cheap to deploy and hard to audit, because you cannot inspect a score the way you can with an external verifier.

Why difficulty is the controlling variable

Snell et al. report that the best strategy depends on prompt difficulty. The intuition is mechanical. On an easy prompt the model’s first sample is usually right, so extra compute buys little. On a prompt within reach, parallel sampling plus a verifier raises the chance of hitting a correct path. On a hard prompt where the model rarely samples a correct chain, voting and best-of-N find nothing to select, and what helps instead is revising or reasoning sequentially, or in the extreme, nothing helps because the model lacks the capability.

This is why “just set N to 64” is poor engineering. A fixed N overspends on easy prompts and underserves hard ones. A controller that estimates difficulty, for example from the agreement rate among the first few samples, can stop early when consensus is clear and escalate only when it is not.

Two selectors compared: ORM and PRM

An outcome reward model (ORM) scores a complete solution, usually trained on whether the final answer was correct. A process reward model (PRM) scores each step. Lightman et al. found step-level supervision substantially better than outcome supervision on MATH when used to rank solutions, and attribute part of the advantage to credit assignment: an ORM sees only that the answer is wrong, while a PRM can point to the step where the chain went off the rails.

Parallel best-of-N pipeline with a generator sampling N chains scored by an ORM or PRM and a majority vote merged into a weighted pick

Figure 2: Best-of-N with two selectors. Final-answer votes and verifier scores can be combined into a weighted pick.

Figure 2 shows both selectors running over the same pool. They are complementary. Voting uses agreement between chains, which is a signal the verifier does not see, while the verifier uses chain quality, which voting ignores. Weighted voting, where each chain’s vote is scaled by its verifier score, is a common hybrid and is the version I would implement first.

A PRM has a cost that an ORM does not: step-level labels. PRM800K needed 800,000 human step labels. Later work has explored generating process labels automatically, but I have not verified any specific numbers here, so treat automated-label claims in vendor material with care and check the paper before relying on them.

Deeper Analysis: Search, Provider Budgets and Cost per Correct Answer

Best-of-N scores complete chains, so it wastes compute finishing chains that were doomed at step two. A PRM lets you score partial chains, which turns generation into search. Snell et al. studied searching against process-based verifier reward models, including beam search and lookahead variants, alongside best-of-N. The idea is straightforward: expand several candidate next steps, score each partial chain, keep the best few and discard the rest.

PRM guided tree search expanding high scoring steps and pruning low scoring branches toward a complete solution

Figure 3: Process reward model guided search prunes weak branches early instead of finishing every chain.

In Figure 3, branches with low PRM scores are pruned after one or two steps and the budget concentrates on the promising subtree. Monte Carlo tree search (MCTS) generalises this by estimating step value through rollouts, trading extra generation for better value estimates. MCTS-style methods carry a heavier engineering burden: you must define what a step is, keep a branching state, and handle the verifier’s calibration on partial reasoning, which is usually worse than on complete reasoning.

Two honest caveats apply. First, search against an imperfect PRM can amplify its errors. The more aggressively you optimise against a learned score, the more you select for chains that fool the scorer rather than chains that are right. Second, step segmentation is task-specific. Splitting on newlines works for tidy math, and fails for code, prose or tool-using agents.

What the provider APIs actually expose

Hosted reasoning models mostly hide the search and expose a single budget dial. The details differ and they change, so I checked the vendor docs for this post. Treat the following as a snapshot from 2026-10-09 and recheck before you ship.

Anthropic. Under the older extended thinking mode you set thinking: {type: "enabled", budget_tokens: N}. The docs state that budget_tokens must be at least 1,024 and less than max_tokens, because thinking counts toward the turn’s output limit (interleaved thinking is the documented exception). The budget is a target, not a cap: the model may stop well short of it, and max_tokens is the hard ceiling. Thinking tokens are billed as output tokens, and the usage object reports how many were reasoning. The docs suggest starting near 1,024 for simple tasks and 16,000 or more for complex ones, and recommend batch processing above 32,000 to avoid timeouts.

The same docs describe a migration. On Claude 4.6 models the manual budget is deprecated but still works, and on Claude 4.7 and later type: "enabled" is rejected with a 400 error. The replacement is adaptive thinking, thinking: {type: "adaptive"}, with depth steered by output_config.effort, where high is the default. With adaptive thinking the model decides whether and how much to think per request, and at lower effort it may skip thinking on easy inputs. In my terms, Anthropic moved the controller from your code into the model.

OpenAI. The reasoning guide describes a reasoning.effort parameter whose supported values are model-dependent and can include none, minimal, low, medium, high, xhigh and max, with defaults that vary by model (several documented models default to medium). Reasoning tokens are not visible through the API but occupy the context window and are billed as output tokens, with the count reported in output_tokens_details.reasoning_tokens. The max_output_tokens limit covers reasoning and visible output together, so a response can come back incomplete before any visible text, leaving you paying for reasoning with no answer. OpenAI recommends reserving at least 25,000 tokens for reasoning and output when you begin experimenting. Raw reasoning is not exposed; you can request a summary.

Google. The Gemini thinking page I fetched documents a thinking_level control (values such as minimal, low, medium and high, varying by model) rather than a numeric budget, with dynamic thinking by default. Thinking tokens are billed, max_output_tokens includes them, and the page advises lowering the thinking level rather than shrinking the token cap to cut cost or latency. Older Gemini 2.5 documentation described numeric thinking budgets; I did not re-verify those figures this run, so I leave them out.

The common pattern across all three: you pay for hidden tokens, the output cap and the reasoning share one pool, and the vendors are converging on coarse levels over precise token counts. For the product-level story, the o3 reasoning analysis goes deeper.

Worked arithmetic: cost per correct answer

Accuracy per dollar, not accuracy alone, is the metric that makes these techniques comparable. The right unit is cost per correct answer: expected spend divided by the probability of being right.

All numbers in this subsection are illustrative, chosen to show the mechanics, not measured on any real model. Assume a model answers a task correctly with probability p = 0.40 on a single sample, a sample costs $0.02 (about 2,000 output tokens at an assumed $10 per million tokens), and wrong answers scatter across ten distinct values rather than concentrating on one.

Strategy Samples Accuracy (illustrative) Spend Cost per correct answer
Single pass 1 0.40 $0.02 $0.050
Majority vote 5 about 0.64 $0.10 about $0.156
Majority vote 9 about 0.80 $0.18 about $0.225
Majority vote 17 about 0.95 $0.34 about $0.36
Majority vote 33 about 0.996 $0.66 about $0.663
Best-of-N, perfect verifier 4 0.870 $0.08 $0.092
Best-of-N, perfect verifier 8 0.983 $0.16 $0.163
Best-of-N, perfect verifier 16 0.9997 $0.32 $0.32

The voting rows come from a Monte Carlo simulation of that scatter assumption; the perfect-verifier rows are the closed form 1 minus (1 minus p) to the N, which is the coverage. Three lessons follow.

First, cost per correct answer always rises with N, even as accuracy rises. Going from 0.40 to 0.80 accuracy with voting multiplies the per-correct cost by 4.5. You are buying reliability, not efficiency, so the question is whether a wrong answer costs more than the premium.

Second, a good verifier is worth a lot. In this toy, 4 samples with a perfect verifier beat 9 samples of voting at less than half the spend. Real verifiers are imperfect, so the true line sits between the two sets of rows.

Third, the wrong-answer distribution matters enormously. If errors collapse onto one popular wrong answer, voting converges on the wrong answer with high confidence. My scatter assumption is generous. Check your own data before trusting any vote.

A runnable best-of-N and majority-vote sketch

The following sketch implements weighted voting with early stopping. It assumes you supply two callables: sample(prompt) returning a chain and its extracted answer, and score(prompt, chain) returning a verifier score in 0 to 1. Everything else is standard library.

from collections import defaultdict
from dataclasses import dataclass
from typing import Callable

@dataclass
class Candidate:
    answer: str
    chain: str
    tokens: int
    score: float = 0.0

def solve(
    prompt: str,
    sample: Callable[[str], Candidate],
    score: Callable[[str, str], float] | None = None,
    n_min: int = 3,
    n_max: int = 16,
    agree_frac: float = 0.75,
    token_budget: int = 40_000,
):
    """Adaptive best-of-N with weighted majority vote and early stop."""
    pool: list[Candidate] = []
    spent = 0
    while len(pool) < n_max and spent < token_budget:
        c = sample(prompt)
        c.score = score(prompt, c.chain) if score else 1.0
        pool.append(c)
        spent += c.tokens
        if len(pool) >= n_min:
            weights = defaultdict(float)
            for p in pool:
                weights[p.answer] += p.score
            best, w = max(weights.items(), key=lambda kv: kv[1])
            total = sum(weights.values()) or 1.0
            if w / total >= agree_frac:      # clear consensus: stop early
                return best, len(pool), spent
    weights = defaultdict(float)
    for p in pool:
        weights[p.answer] += p.score
    best = max(weights.items(), key=lambda kv: kv[1])[0]
    return best, len(pool), spent

The design choices are deliberate. The early-stop rule is the controller from Figure 1 in miniature: easy prompts stop at three samples, hard prompts run to the cap, and the token budget bounds the worst case. Weighting by verifier score is the hybrid from Figure 2. Passing score=None degrades gracefully to plain self-consistency.

Note what the sketch does not do. It does not normalise answers, so “42” and “42.0” count as different votes; add a canonicaliser per task. It also does not handle ties beyond picking the first maximum, and it assumes samples are independent, which fails if you reuse a cached prefix with a fixed seed.

Sequence diagram of a gateway sampling chains from a model, scoring them with a verifier, stopping early on consensus and reporting spend

Figure 4: A gateway that enforces a per-request budget, calls the verifier, stops early and returns a spend report.

Figure 4 puts the sketch inside a service. Putting the policy in a gateway rather than in application code gives you one place to enforce budgets, log spend per request and change N without a deploy. The spend report returned to the application is what lets you compute cost per correct answer on real traffic later.

Trade-offs, Gotchas, and What Goes Wrong

Diminishing returns are the default. Coverage may keep rising with samples, as Large Language Monkeys shows, but the selector does not. The same paper reports that majority voting and reward models plateau beyond several hundred samples. In practice the useful range for voting on a typical task is far smaller, and the marginal cost of each extra sample is constant while its marginal gain shrinks. Plot accuracy against spend on your own evaluation set and look for the knee.

Verifier over-optimisation. A learned ORM or PRM is a proxy. Select the maximum of many samples against it and you increasingly select its blind spots: confident, fluent, wrong chains. The remedy is to hold out an independent check, such as unit tests, a symbolic checker or a second verifier from a different family, and to keep N moderate when the verifier is the only judge. Where you have an automatic, trustworthy verifier, as in code with tests or formal proofs, repeated sampling is at its strongest, which is exactly where Brown et al. saw the biggest gains.

Overthinking. More sequential reasoning is not monotonically better. Long chains can talk themselves out of correct early answers, waste budget re-verifying trivial steps, or loop. I have not pinned the magnitude of this effect to a specific benchmark here, so I state it as a qualitative risk: measure accuracy at several budgets rather than assuming the largest is best. The s1 paper’s own results show the gain from budget forcing is bounded; its AIME24 improvement from 50 to 57 percent is real and also modest.

Truncation you still pay for. Because reasoning tokens share the output cap, a budget set too low can end a response mid-thought. OpenAI documents that a response can be incomplete with no visible output while you are billed for the reasoning. Gemini documents the same behaviour. Set the cap above the reasoning budget with headroom, and alert on incomplete statuses.

Latency stacks. Parallel sampling is cheap in latency only if your serving stack has spare throughput. Under load, N parallel chains queue behind each other, and tail latency is set by the slowest chain. Sequential budgets add latency directly: at 50 tokens per second, a 16,000-token reasoning pass takes about 320 seconds, so user-facing flows often need streaming, background execution or a lower effort setting. That arithmetic is illustrative, since decoding speed varies widely by model and hardware.

Caching and reproducibility. Changing a thinking budget between requests can invalidate prompt-cache breakpoints, according to Anthropic’s docs, so hold the budget stable across a cached conversation. Sampling at temperature above zero also means two runs of the same request can disagree, which complicates regression testing. Fix seeds where the API allows and evaluate on distributions, not single runs.

Hidden reasoning is not an audit trail. The vendors return summaries, not raw chains. Do not build compliance logic that assumes the summary faithfully reflects the computation. If you must reason over sensitive prompts at scale, consider where the extra tokens are processed; see our note on confidential AI inference with TEEs and GPU confidential computing.

It is not a substitute for capability. Snell et al. limit their claim to problems where the base model has non-trivial success. On problems far beyond the model, extra test-time compute cannot rescue it, and a larger or better-trained model is the right move. The 14x figure is a FLOPs-matched comparison under specific conditions in one paper, not a law.

Cost surprises from agents. Multi-step agents multiply the effect: each tool-use turn may carry its own reasoning budget, and prior thinking blocks may be re-billed as input in later turns on some models, according to Anthropic’s migration notes. A modest per-call budget can compound across a twenty-step run. Meter spend per task, not per call.

Practical Recommendations

Start from the cheapest setting that meets your accuracy target and move up only when a measurement says to. Teams that start high and never look back are the ones surprised by invoices.

First, build an evaluation set of at least a few hundred real prompts with known-good answers, and bucket them by difficulty. Without difficulty buckets you cannot see the pattern Snell et al. describe, that different strategies win in different regions.

Second, measure the single-sample baseline and the coverage at N of 4, 8 and 16. If coverage barely exceeds the baseline, parallel sampling will not help and you should look at sequential reasoning or a stronger model. If coverage is high but your selector lags far behind it, invest in the verifier, not in more samples.

Third, prefer verifiers you can trust: tests, schemas, calculators, retrieval checks. Use a learned PRM or ORM when no programmatic check exists, and keep a human-labelled audit slice to catch drift.

Fourth, add the controller. Early stopping on consensus is the cheapest form of compute-optimal allocation and requires no extra model. For local or on-prem deployments where tokens are priced in hardware time rather than dollars, see our guide to running models on a local LLM inference workstation, where the same arithmetic applies with watts and wall-clock in place of price.

Fifth, track cost per correct answer and p95 latency as first-class metrics next to accuracy.

Use this checklist before shipping:

  • A baseline single-sample accuracy and a coverage curve exist for your task.
  • Max output tokens exceed the reasoning budget with headroom, and incomplete responses raise alerts.
  • Answers are canonicalised before voting.
  • An independent check guards against verifier over-optimisation.
  • Per-request spend is logged, and cost per correct answer is reported weekly.
  • Budgets are re-evaluated whenever the vendor changes model versions or defaults.

Decision matrix

Situation Best first move Why Watch out for
Short canonical answers, no verifier Self-consistency with early stop Cheap, no extra model Correlated wrong answers
Code with unit tests Repeated sampling plus tests Verifier is exact, coverage scales Test flakiness, sandbox cost
Math or logic, step structure clear Best-of-N with PRM, or PRM-guided search Step credit assignment Verifier over-optimisation
Open-ended writing or analysis Moderate effort level, single pass No reliable selector Paying for thinking that adds little
Mixed difficulty at scale Difficulty-aware controller Matches Snell et al. finding Controller misrouting hard prompts
Problems beyond model capability Stronger model Compute cannot add capability Wasted spend
Latency-critical user flow Low effort or small budget Sequential cost is linear in tokens Truncation, quality drop

Frequently Asked Questions

What is test-time compute scaling?

Test-time compute scaling means improving a model’s answer by spending more computation at inference rather than by training a bigger model. The spending takes the form of longer reasoning chains, multiple sampled answers combined by voting or a verifier, or search over partial solutions. The model weights stay frozen. Snell et al. (2024) showed that allocating this compute adaptively per prompt can be more effective than scaling parameters on suitable problems.

Is best-of-N better than majority voting?

It depends on whether you have a trustworthy selector. Majority voting needs no extra model but only works when final answers can be compared exactly. Best-of-N with a verifier works on open-ended outputs and can beat voting when the verifier is good, as with unit tests for code. With a weak learned verifier it can pick fluent but wrong chains. Weighted voting, combining both signals, is a sensible default.

What is a process reward model?

A process reward model scores each intermediate reasoning step rather than only the final answer. Lightman et al. trained one on PRM800K, 800,000 human step-level labels, and found process supervision outperformed outcome supervision on MATH, with their best model solving 78 percent of a representative test subset. Its step scores let you rank complete solutions or prune weak branches during tree search.

How do I set a reasoning budget?

Start low and measure. On Anthropic’s older manual mode the minimum budget_tokens is 1,024 and it must be below max_tokens, with 16,000 or more suggested for complex tasks. Newer models use adaptive thinking steered by an effort setting. OpenAI and Google use effort or thinking levels. Always keep the output cap above the budget so responses are not truncated, and compare accuracy at several settings on your own prompts.

Does more thinking always improve accuracy?

No. Gains flatten as budgets grow, long chains can overthink or loop, and the benefit is largest on problems the model can already partly solve. The s1 paper, for example, reports AIME24 accuracy rising from 50 to 57 percent with budget forcing, a real but bounded gain. Extra compute also cannot add knowledge or capability the model lacks, so a stronger model sometimes wins.

How much does test-time scaling cost?

Cost scales with total tokens generated, including hidden reasoning tokens that providers bill as output. Parallel sampling multiplies cost by N, sequential reasoning by chain length. The useful metric is cost per correct answer. In the illustrative example above, raising accuracy from 0.40 to about 0.80 by voting raised cost per correct answer from $0.05 to about $0.225. Measure your own numbers before budgeting.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *