Next-Token Sampling, Explained

After reading this you will be able to read a language model's next-token distribution, predict exactly how temperature, top-k and top-p reshape it, and choose decoding settings on purpose instead of by trial and error.

What next-token sampling actually is

A language model does not "decide" what word comes next. It outputs one number per candidate token, called a logit, and those numbers become a probability distribution over the whole vocabulary. Decoding is the step that picks one token from that distribution. Every setting an LLM API exposes, temperature, top-k, top-p, repetition penalty, is arithmetic applied to that single distribution before you draw a sample.

Here is the hook. Suppose the model has read the prompt the cat sat on the and the top three candidates carry logits 3.0 for mat, 1.0 for rug, and 0.0 for floor. After the softmax those become probabilities of about 0.844, 0.114 and 0.042. Greedy decoding always prints mat. Raise the temperature to 2.0 and the same three logits become 0.506, 0.307, 0.187, so floor now appears almost one time in five. The model never changed. Only the arithmetic on its output changed.

The playground trains a small word-level n-gram model in your browser, shows you the real distribution as bars, and lets you reshape it live. The model is tiny, but the sampling pipeline downstream of the logits is the same one that GPT-class APIs run.

When this view helps and when it misleads

Use this mental model whenever you tune decoding: picking between deterministic output and varied output, debugging a model that loops, or deciding what temperature and top_p to send to an API. The distribution picture explains all of it.

Be careful about one thing. An n-gram model conditions only on the last one to three words, then backs off to shorter context when it has not seen the full one. A transformer conditions on thousands of tokens. So the content of the distribution here is far cruder than a production model's, and long-range coherence is not something a toy can show. What transfers exactly is the reshaping math. A temperature of 0.7 does the identical thing to a 15-token bar chart and to a 50,000-token vocabulary.

If you want to see where those tokens come from in the first place, the LLM Tokenizer Visualizer shows how text splits into the units a model actually scores.

The softmax and temperature

The distribution comes from the softmax of the logits. For token i with logit z_i and temperature T:

p_i = \frac{e^{z_i / T}}{\sum_{j} e^{z_j / T}}

Here z_i is the raw score for token i, the sum in the denominator runs over every candidate j, and e is Euler's number, about 2.718. Dividing every logit by T before exponentiating is the whole trick.

When T = 1 you get the model's native distribution. When T \lt 1 the gaps between logits grow, so the largest logit dominates and the distribution sharpens toward its mode. When T \gt 1 the gaps shrink and the distribution flattens toward uniform. As T \to 0 the largest logit wins with probability approaching 1, which is exactly greedy argmax.

A useful summary number is entropy, the average surprise of the distribution in bits:

H = -\sum_{i} p_i \log_2 p_i

Entropy is 0 when one token has all the mass and reaches its maximum, \log_2 n bits, when all n tokens are equally likely. The playground shows this readout so you can watch a slider trade certainty for variety as a single number.

A worked example you can reproduce

The demo trains on the built-in corpus and starts from the field defaults. Take the three logits from the hook, 3.0, 1.0, 0.0 for mat, rug, floor, and work the softmax by hand at three temperatures.

Softmax at three temperatures

  1. At T = 1, exponentiate the logits: e^3 = 20.09, e^1 = 2.718, e^0 = 1.0. The sum is 23.80.
  2. Divide each by the sum: 20.09 / 23.80 = 0.844, 2.718 / 23.80 = 0.114, 1.0 / 23.80 = 0.042.
  3. At T = 0.5, divide logits by 0.5 first: 6, 2, 0. Then e^6 = 403.4, e^2 = 7.389, e^0 = 1.0, sum 411.8. Probabilities: 0.980, 0.0179, 0.0024.
  4. At T = 2, divide logits by 2: 1.5, 0.5, 0. Then e^1.5 = 4.482, e^0.5 = 1.649, e^0 = 1.0, sum 7.131. Probabilities: 0.629, 0.231, 0.140.

The entropy at T = 1 is -(0.844 \log_2 0.844 + 0.114 \log_2 0.114 + 0.042 \log_2 0.042) = 0.752 bits. At T = 0.5 it drops to 0.152 bits, nearly certain. At T = 2 it rises to 1.30 bits, close to the 1.585-bit maximum for three equal outcomes.

Lower temperature stacks mass on the top token; higher temperature spreads it toward the tail. The logits never move.

With logits 3.0, 1.0, 0.0 the softmax gives probabilities 0.844, 0.114, 0.042 at temperature 1.0; at temperature 0.5 it gives 0.980, 0.018, 0.002; at temperature 2.0 it gives 0.629, 0.231, 0.140. Lower temperature sharpens toward the top token, higher temperature flattens toward equal thirds.

Truncation: top-k and top-p

Temperature reweights every token but never removes one. Truncation removes tokens outright, then renormalizes what survives so the kept probabilities sum back to 1.

Top-k keeps the k most probable tokens and zeroes the rest. With k = 2 applied to the T = 1 distribution above, you keep mat at 0.844 and rug at 0.114, drop floor, and renormalize by the surviving sum 0.958: mat becomes 0.881 and rug becomes 0.119. floor can never be sampled no matter how high you push the temperature afterward.

Top-p (nucleus sampling) sorts tokens by probability and keeps the smallest set whose cumulative probability reaches p. With p = 0.9 you add mat (0.844), then rug to reach 0.958, which crosses 0.9, so you stop and drop floor. Same two survivors here, renormalized to 0.881 and 0.119.

The difference shows up when the distribution's shape varies. Top-k always keeps exactly k tokens even when the model is unsure and many are plausible. Top-p keeps few tokens when the model is confident and many when it is uncertain, so it adapts to the local entropy. That is why p = 0.9 is a common default while a fixed k often needs retuning per context.

Order matters. The playground applies temperature first, then truncation. Raising temperature after top-p would let cut tokens back in, so APIs fix the order: reshape with temperature, then cut, then renormalize, then sample.

Reading the bars and the repetition penalty

In the playground the gray outlines are the raw model probabilities, the filled bars are the distribution you will actually sample from, and grayed-out bars are candidates cut by truncation. When a filled bar sits taller than its gray outline, renormalization lifted it because tokens below it were removed. Watch the entropy readout as you drag: a low number means near-deterministic output, a high number means the next token is close to a coin toss.

The repetition penalty divides the probability (or logit) of any token seen recently by a factor greater than 1. Suppose mat already appeared and the penalty is 1.3. Its probability 0.844 becomes 0.844 / 1.3 = 0.649 before renormalizing, which pulls mass toward rug and floor. This is the direct fix for the classic failure mode below.

Common mistakes

Greedy decoding with no repetition penalty loops. The argmax at T \to 0 is deterministic, so once the model enters a state whose top token leads back to that same state, it repeats forever: "the the the" or a two-word cycle. Raising temperature or adding a penalty above 1.0 breaks the cycle.

Three more errors are worth naming. First, stacking a high temperature with a wide top-p produces rambling, because you both flatten the distribution and refuse to cut its tail. Second, setting k = 1 is just greedy decoding, so temperature has no effect at all once one token survives. Third, treating temperature as a "creativity" dial is misleading: it only rescales existing logits. If the model assigns a good next word almost no probability, no temperature will conjure a good distribution, it will only amplify whatever mediocre tail the model already has.

Related tools

Once you understand the distribution, two nearby tools cover the steps around it. The LLM Tokenizer Visualizer shows how raw text becomes the tokens the model scores, which is the input side of this same pipeline. When you move from experimenting to running real prompts at real settings, the LLM API Cost & Context Planner estimates spend per day and month, including context length and prompt caching.

Frequently asked questions

What temperature should I use?

For factual or code tasks where you want the most likely answer, use 0.0 to 0.3. For varied prose, 0.7 to 1.0 is typical. Above 1.3 the tail grows fast and output gets noisy. There is no universal best value: it depends on how peaked your model's distributions already are, which the entropy readout lets you check.

Should I use top-k or top-p?

Prefer top-p (nucleus) as a default around 0.9 to 0.95, because it adapts to how confident the model is at each step. Use top-k when you want a hard cap on how many tokens can ever be considered. Many APIs let you set both, and they compose: the effective candidate set is the intersection.

Why does temperature 0 give the same answer every time?

At T \to 0 the softmax puts probability approaching 1 on the single largest logit, so sampling always draws that token. There is no randomness left to draw from. This is greedy argmax decoding, and it is deterministic given the same prompt.

Does an n-gram model tell me anything about GPT?

The content does not transfer, because an n-gram sees only the last few words while a transformer sees thousands. The reshaping math transfers exactly. Temperature, top-k, top-p and the repetition penalty operate on a probability distribution the same way regardless of how that distribution was produced.

What does the entropy readout mean in practice?

It is the average number of bits needed to encode the next token under the current distribution. Near 0 bits means the output is effectively fixed. For three candidates the maximum is 1.585 bits, reached only when all three are equally likely. Watching it fall as you lower temperature is the clearest single signal of how deterministic your settings are.