Softmax temperature
Temperature T controls how spread out a model's next-token probabilities are. Move T and watch the four candidates. The order never changes. Only the spread does.
One decoding step
Context: “The cat sat on the mat.” The model scores four candidate tokens.
What the bits mean
Each token has a surprise of −log₂ p bits. One bit is one fair yes/no question.
Entropy is the average surprise of one draw. A count of the possible tokens says 4 at every T, even at T = 0.25, where cat wins 98% of draws; entropy says about 1.11 choices there.
T = 1 is plain softmax: the model's own distribution.
Divide, exponentiate, normalize
The sampler does three operations on the logits. The table updates with the slider. The last column is the bar chart above, as numbers. Press the 2.0 preset to see the worked example from the prose.
| token | model outputlogit z | step 1z / T | step 2exp(z / T) | step 3 · ÷ sumprobability p |
|---|
Only the gap matters, and T divides the gap
The ratio of two probabilities depends on the gap between their logits, divided by T: the scaled gap. T does not change the gap itself. A small T makes the scaled gap large, so the top token dominates. A large T makes the scaled gap small, so the ratio goes toward 1.
Put another way, every ratio becomes its T = 1 value raised to the power 1/T. At T = 2: (0.609 / 0.224)^(1/2) = 1.65. Adding T to every logit would do nothing, as the toggle above shows; dividing is what scales the gaps.
Try a simpler normalizer
Softmax uses exp because exp turns every logit into a positive weight, keeps the order, and turns a gap into a ratio: e^(a−b) = e^a / e^b. Pick another way to turn the scaled logits into shares, and see what breaks. Move T too.
| normalizer | cat | dog | fox | owl |
|---|
The name comes from physics: the Boltzmann distribution is p ∝ exp(−E / kT). A logit plays the role of −E, and T is the temperature.
What the sampler produces
Draw 60 tokens at the current temperature. At a low T almost every draw is “cat”. At a high T the output varies more.
What temperature does not do
Division by a positive T keeps the order of the logits. Cat is first at every temperature. A negative T would reverse it: at T = −1, owl becomes most likely (0.71). APIs do not allow T below 0.
Division by 0 is undefined. APIs pick the top token every time.
Real vocabularies are much larger. Many APIs also apply top-k or top-p after temperature. The behavior shown here is the same.
It controls spread only. A flatter distribution gives more varied output. It adds no knowledge.
Three facts
- The sampler divides every logit by T before softmax.
- pi / pj = exp((zi − zj) / T). Low T sharpens; high T flattens.
- The order of tokens never changes. T = 0 means greedy; T → ∞ means uniform.