Next-Token Sampling Playground illustration

Next-Token Sampling Playground

Every "creative" or "deterministic" LLM setting is just arithmetic on one probability distribution. This playground trains a small word-level n-gram language model (context of 1–3 words, with backoff) on an editable corpus, right in your browser. Type a prompt and the model’s true next-token distribution appears as a bar chart of the top 15 candidates. Then move the sliders — temperature, top-k, top-p and repetition penalty — and watch the bars reshape in real time: gray outlines show the raw model probabilities, filled bars the distribution you will actually sample from, and grayed-out bars the candidates cut by truncation. Readouts track the entropy of the current distribution. Press Generate to stream tokens one at a time with your settings, or run the greedy vs high-temperature preset side by side to see why one repeats itself and the other rambles.

Runs 100% in your browser — models are trained and computed locally on your device.

Notes

  • Temperature divides the log-probabilities before the softmax: T < 1 sharpens the distribution toward the mode, T > 1 flattens it toward uniform, and T → 0 is exactly greedy argmax decoding.
  • Top-k keeps only the k most probable tokens and top-p keeps the smallest set whose probabilities sum to p — both then renormalize, so cut tokens get probability zero no matter how the temperature reshaped them.
  • The repetition penalty divides the probability of every token that already appeared recently, which is why greedy decoding with no penalty so often falls into loops.
  • An n-gram model only conditions on the last few words, unlike a transformer’s long context — but the sampling pipeline downstream of the logits is exactly the same one production LLM APIs expose.
  • Runs 100% in your browser — models are trained and computed locally on your device.