Kolibri-1 Model Plays Atari Breakout Real-Time via Output Probabilities
Developers demonstrated Aleph Alpha's Kolibri-1 model playing Atari Breakout in real-time without any fine-tuning. By bypassing text generation and directly mapping the model's raw token output probabilities to four discrete game actions, they achieved an inference latency under 25 milliseconds per move. This experiment demonstrates that large language models can act as ultra-low-latency controllers by leveraging raw output probabilities rather than full autoregressive text generation. It highlights practical possibilities for open-weight models in real-time agent control, interactive applications, and robotics. The setup avoids autoregressive token generation entirely, extracting logits from the model's final softmax layer to instantly select among four predefined actions. Because Kolibri-1 uses a Mixture-of-Experts (MoE) architecture that activates only 3.46 billion parameters out of 78.1 billion per token, optimized probability extraction enables extreme speed.
## BACKGROUND
Standard Large Language Model (LLM) text generation uses autoregressive decoding, where tokens are predicted sequentially, adding latency at every word. Before selecting a output token, the LLM's final linear and softmax layers compute raw probability distributions across its whole vocabulary. By reading these action-token probabilities directly without initiating multi-token text generation, applications can achieve near-instantaneous decision-making cycles.