01Dynamic VRAM Swapping Boosts MoE Prefill Speeds on Dual Consumer GPUsREDDIT · /u/Extension-Bid-639 · reddit.com · Sep 10, 02:39 AM11h
02Qwen3.8-Flash-Next Benchmark: SGLang Beats llama.cpp by 7.3x in Time-to-First-TokenREDDIT · /u/FantasticNature7590 · reddit.com · Sep 08, 07:26 PM1d
03LayerStoRm Enables Fast MoE Model Streaming Beyond GPU VRAM LimitsREDDIT · /u/CharacterBumblebee99 · reddit.com · Sep 07, 01:14 AM3d
04New Tool "cache-pressure" Validates Local LLM KV Cache Eviction under PressureREDDIT · /u/t4a8945 · reddit.com · Sep 06, 09:38 AM4d
05SGLang v0.5.19 Released with New Model Support and Performance OptimizationsGITHUB · Qiaolin-Yu · github.com · Sep 05, 02:27 AM5d
06MTP Merged into ik_llama.cpp, Doubling Qwen Inference Speed for CodingREDDIT · /u/Alternative_Will5974 · reddit.com · Sep 03, 04:30 PM6d
07Perplexity Open-Sources 'Lily' Inference Engine for Apple SiliconREDDIT · /u/Specter_Origin · reddit.com · Sep 02, 10:13 PM7d
08Why vLLM Achieves Drastically Faster Prefill Speeds Than llama.cpp on Multi-GPU SystemsREDDIT · /u/dangerous_inference · reddit.com · Sep 01, 05:05 PM8d
09Developer Pushes 27B LLM Inference to 2,000 Prefill Tokens/Sec on RTX 3090REDDIT · /u/iamMess · reddit.com · Sep 01, 11:43 AM9d
10ExLlamaV3 Update Adds MoE CPU Offloading and Self-Calibrated QuantizationREDDIT · /u/Unstable_Llama · reddit.com · Sep 01, 07:14 AM9d
11llama.cpp v0.3.0 Released with GLM-4.5-Air and DeepSeek 4 SupportGITHUB · github-actions[bot] · github.com · Aug 25, 10:22 AM16d
12Moore Threads Releases Prefill-as-a-Service Whitepaper for MTT S5000 GPUsRSS · IT HOME · ithome.com · Aug 24, 02:09 AM17d
13The Maturation and Adoption of Speculative Decoding in 2026 LLM InferenceREDDIT · /u/Ok-River5924 · reddit.com · Aug 10, 08:02 AM31d
14Developer Runs LFM-2.5 2.6B Model on Mobile CPU at 17 Tokens/SecondREDDIT · /u/trikboomie · reddit.com · Aug 05, 02:20 PM36d
15llama.cpp b10228 Released with Support for DeepSeek-V4 MTP and DSparkGITHUB · github-actions[bot] · github.com · Aug 02, 01:28 PM39d
16llama.cpp Adds MTP and DSpark Support for DeepSeek V4 FlashREDDIT · /u/rmhubbert · reddit.com · Aug 02, 12:58 PM39d
17OpenAI Uses GPT-5.6 Sol to Optimize Inference and Cut Serving Costs by 20%RSS · IT HOME · ithome.com · Jul 30, 06:06 AM42d
18llama.cpp Releases Build b10150 with Backend Offloading AdjustmentsGITHUB · github-actions[bot] · github.com · Jul 27, 12:45 PM45d