Tag: EXL3
-
New TensorFold LLM Model Serving Software Dramatically Improves Decode Speed on Macs and Sparks
by Chris DePuy, 650 Group / September 27, 2026 Two repositories shipped from one maintainer this week and made the same promise in two different places. On September 25 and 26, a developer publishing under the name Ash Hart released TensorFold, a local inference engine for Apple Silicon and NVIDIA GPUs, and Imprint, a tool…
-
GLM-5.3-Flash on a Single DGX Spark: Three Community Quantizations Fit Where the Vendor’s Does Not
by Chris DePuy / September 13, 2026 On September 9, NVIDIA published an official NVFP4 quantization of Z.ai’s GLM-5.3-Flash. Its Model Optimizer recipe leaves the key-value cache at FP8 and compresses only the linear layers inside the transformer blocks — the sparse-MoE shared experts and the dense MLP — from 16 bits down to 4,…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
