21llama.cpp Releases Build b10736 with Fix for Architecture Test LoggingGITHUB · github-actions[bot] · github.com · Sep 01, 01:29 PM2d
22llama.cpp Adds Fixes and Multi-Token Prediction Support for Qwen ModelsREDDIT · /u/jacek2023 · reddit.com · Sep 01, 08:28 AM3d
23llama.cpp b10729 Adds Metal Flash-Attention Vector Tunings for M1 UltraGITHUB · github-actions[bot] · github.com · Aug 31, 10:22 PM3d
24llama.cpp Release b10721 Fixes WebGPU Backend CrashesGITHUB · github-actions[bot] · github.com · Aug 31, 03:54 PM3d
25llama.cpp Release b10711 Fixes Hexagon DSP Backend BugGITHUB · github-actions[bot] · github.com · Aug 31, 04:45 AM4d
26llama.cpp Releases Build b10708 with GGML Backend Bug FixGITHUB · github-actions[bot] · github.com · Aug 31, 03:25 AM4d
27llama.cpp b10707 Optimizes KV-Cell Scan for Faster Long-Context InferenceGITHUB · github-actions[bot] · github.com · Aug 31, 03:03 AM4d
28llama.cpp Release b10700 Renames Lazy Mode CLI ArgumentGITHUB · github-actions[bot] · github.com · Aug 30, 06:33 PM4d
29llama.cpp b10694 Released with RPC Fix for Older macOS VersionsGITHUB · github-actions[bot] · github.com · Aug 30, 01:06 PM4d
30llama.cpp b10693 Adds Qualcomm Hexagon NPU Device Discovery and Lazy AllocationGITHUB · github-actions[bot] · github.com · Aug 30, 12:43 PM5d
31llama.cpp Release b10691 Fixes Metal Pipeline Crash on Apple SiliconGITHUB · github-actions[bot] · github.com · Aug 30, 11:58 AM5d
32llama.cpp b10687 Optimizes OpenCL for Adreno GPUsGITHUB · github-actions[bot] · github.com · Aug 29, 06:23 PM5d
33llama.cpp Release b10662 Introduces Unified KV Cache Configuration ArgumentGITHUB · github-actions[bot] · github.com · Aug 27, 09:07 PM7d
34llama.cpp Merges Support for Qwen3.8-Flash-Next ModelREDDIT · /u/jacek2023 · reddit.com · Aug 27, 07:34 PM7d
35llama.cpp Release b10648 Simplifies MiniMax-01 Model GraphGITHUB · github-actions[bot] · github.com · Aug 27, 11:59 AM8d
36llama.cpp b10645 Released with CPU Offloading for Dense FFN WeightsGITHUB · github-actions[bot] · github.com · Aug 27, 10:02 AM8d
37llama.cpp b10643 Released with Multi-NPU Hexagon SupportGITHUB · github-actions[bot] · github.com · Aug 27, 02:14 AM8d
38llama.cpp Release b10642 Introduces Token ID Tracking to KV CellGITHUB · github-actions[bot] · github.com · Aug 26, 10:00 PM8d
39llama.cpp Releases Build b10639 with Vulkan Warp Size FixGITHUB · github-actions[bot] · github.com · Aug 26, 04:51 PM8d
40llama.cpp Release b10631 Updates Tensor Initialization in ggml-metaGITHUB · github-actions[bot] · github.com · Aug 26, 06:04 AM9d