~/LOCAL LLM/local-ai-practitioner-achieves-3x-speedup-with-ornith-1-5-35b-a3b

Local AI Practitioner Achieves 3x Speedup with Ornith 1.5 35B-A3B Model

A practitioner running local LLMs on a dual RTX 5070 Ti setup achieved a 3x inference speed increase—reaching around 180 tokens per second—by switching from Qwen3.8 27B to Ornith 1.5 35B-A3B. The new model maintained 100% pass rates across the user's custom multi-step agent and coding evaluation benchmarks. This real-world benchmark demonstrates how modern Mixture-of-Experts architectures and hybrid linear attention can drastically improve throughput on consumer hardware without degrading functional capabilities. It highlights that fast generation combined with long context windows (128K+) is becoming practical for desktop AI agent workflows. Ornith 1.5 35B-A3B achieves its high inference speed because only about 3B parameters are active per token out of 256 total experts, and only 10 of its 41 layers use full attention. Consequently, the KV cache grows very slowly, adding only about 2GB of RAM overhead when increasing the context window from 128K to 256K tokens.

## BACKGROUND

Running AI models locally requires balancing GPU memory (VRAM) constraints against inference speed and context length. Mixture-of-Experts (MoE) models reduce compute overhead per token by dynamically routing inputs to a small active subset of specialized parameters ('experts'). Additionally, hardware extensions like OCuLink enable desktop setups to seamlessly attach external secondary GPUs for multi-GPU inference.

## REFERENCES

## KEYWORDS

#Local LLM#Inference Optimization#Hardware Benchmarks#AI Agents#Model Evaluation

$ subscribe --daily

Local AI Practitioner Achieves 3x Speedup with Ornith 1.5 35B-A3B Model | Daily News