llama.cpp b11472 Optimizes Zero-Temperature Sampling with Greedy Selection
Release b11472 of llama.cpp introduces greedy token selection for eligible zero-temperature sampling chains across CPU execution, grammar-constrained, and reasoning-budget paths. It simplifies zero-temperature logic by bypassing full distribution sampling when temperature is zero or after top-k filtering reduces candidates to a single token. This update streamlines zero-temperature inference, improving execution efficiency and code maintainability during deterministic decoding. It ensures that specialized paths, such as grammar-constrained structured outputs, share the same optimized greedy sampling behavior as standard text generation. Distribution sampling remains active whenever dynamic temperature or explicit token probabilities are requested by the user. The update also updates sampler behavior tests to cover shared zero-temperature logic across CPU and constrained sampling chains.
## BACKGROUND
Large language models output next-token probabilities, where the temperature parameter controls randomness; setting temperature to zero usually triggers greedy decoding to select the highest-probability token. Grammar-constrained decoding masks invalid tokens at each step to enforce strict format compliance (such as JSON), requiring sampler logic to handle deterministic token selection correctly.