01IFM Releases K2-Horizon-MoVA-36B-A4B GGUF Model with 512K Context WindowREDDIT · /u/jacek2023 · reddit.com · Sep 03, 01:47 PM22h
02KV Cache Growth is Overtaking Parameter Count as the Primary VRAM BottleneckREDDIT · /u/jonejy · reddit.com · Sep 03, 06:04 AM1d
03Community Speculates on Shift in Alibaba's Qwen Model Release StrategyREDDIT · /u/chocofoxy · reddit.com · Sep 02, 02:14 PM1d
04Community Speculates on Qwen 4, Extended Reasoning, and Engram ArchitectureREDDIT · /u/LegacyRemaster · reddit.com · Sep 02, 07:53 AM2d
05LocalLLaMA Community Speculates on Upcoming Open-Source LLM Parameter SizesREDDIT · /u/Porespellar · reddit.com · Sep 01, 06:54 PM2d
06Reddit User Exposes Potentially Mislabeled GGUF Model Quantizations by AtomicChatREDDIT · /u/po_stulate · reddit.com · Sep 01, 03:27 PM2d
07llama.cpp Release b10730 Optimizes Qwen Prompt Processing SpeedGITHUB · github-actions[bot] · github.com · Sep 01, 05:08 AM3d
08Tencent Open-Sources Compressed 770B Hy4 MoE Model, Shrinking Footprint to 214GBRSS · IT HOME · ithome.com · Sep 01, 04:31 AM3d
09iFLYTEK Open-Sources Spark X2.5 Edge LLMs with 1M Token ContextRSS · IT HOME · ithome.com · Sep 01, 03:20 AM3d
10Inquiry on Running Qwen 3.8 Flash Locally Using llama.cppREDDIT · /u/No_Algae1753 · reddit.com · Aug 31, 06:36 PM3d
11Release of Uncensored GGUF Models Including 1M Context LongCat-Flash-Lite-SparseREDDIT · /u/LLMFan46 · reddit.com · Aug 30, 02:16 PM4d
12Tencent Compresses Hy4-Preview Model to 200GB GGUF with 98% Performance RetainedREDDIT · /u/RedditUsr2 · reddit.com · Aug 29, 02:31 PM5d
13Evaluating Qwen 3.8 Flash Next on 4x RTX 3090 GPUsREDDIT · /u/Acceptable_Adagio_91 · reddit.com · Aug 28, 11:38 PM6d
14Tencent Releases Hunyuan Hy4 Preview, a 770B Parameter Open-Source ModelRSS · IT HOME · ithome.com · Aug 28, 06:15 AM7d
15Ant Group Introduces Ling-3.0-flash-Fin Finance Model, Set to Open-Source Next WeekRSS · IT HOME · ithome.com · Aug 28, 02:41 AM7d
16GLM-5.3 Flash GGUF Quantizations Released via Unsloth for Local ExecutionREDDIT · /u/ElementNumber6 · reddit.com · Aug 27, 02:55 PM7d
18Open-Weight Release of GLM-5.3-Flash with Hybrid Sparse-Linear AttentionREDDIT · /u/No_Afternoon_4260 · reddit.com · Aug 26, 03:17 PM8d
19User Claims 27B Model Outperforms Frontier Models in Agentic TasksREDDIT · /u/Gohab2001 · reddit.com · Aug 26, 08:45 AM9d
20Thomson Reuters Releases Thomson-1.0-Small, a Specialized Law and Tax LLMREDDIT · /u/RedditUsr2 · reddit.com · Aug 26, 03:10 AM9d