Liquid AI Releases d1-3B and d1-omni Zero-Output-Token Multimodal Decision Models
Liquid AI announced d1-omni-600M and d1-3B, two edge-sized decision models built on Liquid Foundation Model (LFM) architectures that process text, images, and audio. Unlike generative models, they return typed answers and calibrated probabilities directly from logit distributions in a single forward pass without generating output tokens. By eliminating autoregressive token generation, these models significantly reduce inference latency and computational overhead for real-time decision-making tasks. This approach enables deterministic, ultra-low-latency AI decisions on edge devices without the costs associated with billed LLM output tokens. The 587M-parameter d1-omni handles text, multi-image tiling, and audio clips up to 30 seconds, while the 3B-parameter d1-3B scored 48.57 on Decision Index 0.2.1, beating larger models like Decider 35B-A3B. Both models offer sub-10ms decision latencies on GPUs like the RTX 4090 (8 ms) and AMD MI325X (9 ms), with GGUF format files published on Hugging Face.
## BACKGROUND
Standard generative Large Language Models (LLMs) construct outputs one token at a time through repeated autoregressive decoding loops, which introduces decoding latency and variable output length. Liquid AI's Liquid Foundation Models (LFMs) utilize efficient hybrid architectures, allowing decision variants like d1 to compute probability distributions over structured prompt choices directly in a single forward pass.