~/LLM/llama-cpp-v0-6-0-released-with-extended-batch-api-and-apple

llama.cpp v0.6.0 Released with Extended Batch API and Apple GPU Optimizations

llama.cpp version 0.6.0 introduces the llama_batch_ext extended batch API for mixed token/embedding inputs and state embeddings, alongside support for new architecture families like the 320B GLM-5.3-Flash hybrid model. It also delivers substantial performance gains, including up to ~3x faster matrix multiplication on Apple GPUs and MTP speculative decoding for Qwen4Exp. As a foundational framework for open-source LLM inference, llama.cpp's updates enable developers to efficiently run complex multi-modal and decision models on local hardware. The performance optimizations across Metal and Vulkan significantly lower latency and resource requirements for consumer devices. The release adds a new /v1/systemone server endpoint for decision models and bumps session storage formats to LLAMA_SESSION_VERSION 11. Low-level hardware changes include a Metal tensor-API flash attention kernel for F16 KV caches and sparse flash attention support for quantized KV caches on Vulkan.

## BACKGROUND

llama.cpp is a popular open-source C/C++ library engineered for low-latency, quantized Large Language Model (LLM) inference across diverse hardware. Speculative decoding, such as Multi-Token Prediction (MTP), accelerates inference speed by predicting multiple candidate tokens ahead of time using specialized draft layers before verifying them in parallel with the main model.

## REFERENCES

## KEYWORDS

#llm#llama-cpp#machine-learning#ai-infrastructure#hardware-acceleration

$ subscribe --daily

llama.cpp v0.6.0 Released with Extended Batch API and Apple GPU Optimizations | Daily News