~/LLM/ornith-1-5-model-variants-paired-with-dflash-draft-models-released-on

Ornith-1.5 Model Variants Paired with DFlash Draft Models Released on Hugging Face

New Hugging Face repositories have been released pairing variants of the Ornith-1.5 model—including 9B, 35B-A3B, and 397B parameters—with DFlash draft models. These paired checkpoints are designed to enable speculative decoding for accelerated local LLM inference. Speculative decoding significantly reduces token generation latency without altering the output quality of the primary model. Providing ready-to-use DFlash draft models for Ornith-1.5 models up to 397B parameters makes high-speed local inference accessible for large-scale open models. DFlash uses a block diffusion architecture to generate multiple draft tokens in parallel within a single forward pass, rather than drafting them autoregressively. It also reuses context hidden states from the target model via KV injection, avoiding redundant computation during token drafting.

## BACKGROUND

Large Language Models typically generate text autoregressively, predicting one token per step, which often leaves parallel hardware compute capacity underutilized. Speculative decoding speeds up generation by having a fast draft model propose candidate tokens that the larger target model verifies in parallel. DFlash is a drafting framework that replaces standard sequential draft models with lightweight parallel block diffusion models.

## REFERENCES

## KEYWORDS

#LLM#Speculative Decoding#Inference Optimization#Open Source AI

$ subscribe --daily