Recent Blog Posts
-
GLM-5.3-Flash on a Single DGX Spark: Three Community Quantizations Fit Where the Vendor’s Does Not
by Chris DePuy / September 13, 2026 On September 9, NVIDIA published an official NVFP4 quantization of Z.ai’s GLM-5.3-Flash. Its Model Optimizer recipe leaves the key-value cache at FP8 and compresses only the linear layers inside the transformer blocks — the sparse-MoE shared experts and the dense MLP — from 16 bits down to 4,…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
-
DFlash2 Beats DSpark for Qwen3.8-27B on One DGX Spark — and Jumps to Apple Silicon
by Chris DePuy / August 20, 2026 A week ago the Qwen3.8-27B decode race on a single NVIDIA DGX Spark was a contest between DSpark and MTP. This week the leading recipes on the same hardware have switched to a different engine: DFlash2, the block-diffusion speculative decoder that first shipped with DeepSeek, is now beating…
-
GLM-5.3-Flash Day Two: DFlash 2 Attacks the Decode Path While 1M-Context KV Moves Onto Two DGX Sparks
by Chris DePuy / August 28, 2026 A day after Z.ai opened weights on GLM-5.3-Flash, the desk-side community pushed the model past what I thought was its practical limit on small hardware. Two bottlenecks — the decode path and the KV cache — were attacked in parallel, and both moved in the same 24 hours.…
-
Grok Bot, Two Weeks In: The Always-On Agent Workforce Gets a Shared Computer and a Shared Security Boundary
by Chris DePuy / August 29, 2026 xAI launched Grok Bot in early beta on August 11, and in the two weeks since, the community has made the platform’s ambition concrete faster than the company’s own launch materials do. The product is not another chat interface — a Bot gets its own cloud computer, signs…
-
DeepSeek V4 Flash 0731 on the Multi-Node DGX Spark: From One Pair to an Eight-Node Cluster
by Chris DePuy / August 4, 2026 When DeepSeek shipped the official DeepSeek-V4-Flash-0731 weights on July 31, the community benchmarks that followed concentrated on single and dual DGX Spark units. Over the past several days a more interesting pattern has emerged as the same checkpoint moves up the cluster ladder, from a single GB10 to…
-
Kimi K3 Open at 2.8 Trillion Parameters Closes The Gap
by Chris DePuy / July 28, 2026 On July 27, Moonshot AI made good on its promise: the full weights of Kimi K3 — at 2.8 trillion total parameters the largest open-weight model ever released — went live on Hugging Face under a Modified MIT license. The release includes the model’s technical report and a…
-
Poolside’s Laguna S 2.1 118B Punches Above Its Weight
by Chris DePuy / July 28, 2026 This week, the local AI community has been running Poolside’s latest open-weight model, Laguna S 2.1, on everything from single DGX Sparks to 3090 Quad workstations. Released on July 21, Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, a 1M-token…
-

1T Inkling Model at 1-bit is 86% smaller
AI-generated concept image of model compression via 1-bit quantization. Thinking Machines Lab’s Inkling — a 975B-parameter (41B active) multimodal with 1M context, vision, and audio (HF model card) — launched day one with a dynamic 1-bit GGUF courtesy of Unsloth. The Compression Format Size Vs. Full BF16 (full) 1.9 TB — NVFP4 ~490 GB 74%…
Top 10 Companies
Agentic AI AI Agents Cloudflare DeepSeek DeepSeek V4.1-Flash Deskside AI DFlash DFlash2 DGX Spark DSpark FP8 GB10 GLM-5.3-Flash inference benchmark KV Cache Local-Inference long context MoE MXFP4 Nemotron NVFP4 NVIDIA Open Weights Quantization Qwen SGLang Speculative Decoding Tensor Parallelism Unified Memory vLLM
