Tag: MoE
-
DeepSeek-V4.1-Flash on Four DGX Sparks: The Profile Day One Only Configured Now Boots, Serves and Passes a Million-Token Needle Test
by Chris DePuy / September 14, 2026 Late Sunday evening Pacific time, Mia AI Lab pushed a four-rank serving profile for DeepSeek-V4.1-Flash to its DGX Spark recipe, the same lab whose September 13 post on X put the model on four DGX Sparks and invited readers to run the stack themselves. The significance is not…
-
GLM-5.3-Flash on a Single DGX Spark: Three Community Quantizations Fit Where the Vendor’s Does Not
by Chris DePuy / September 13, 2026 On September 9, NVIDIA published an official NVFP4 quantization of Z.ai’s GLM-5.3-Flash. Its Model Optimizer recipe leaves the key-value cache at FP8 and compresses only the linear layers inside the transformer blocks — the sparse-MoE shared experts and the dense MLP — from 16 bits down to 4,…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
-
GLM-5.3-Flash Day Two: DFlash 2 Attacks the Decode Path While 1M-Context KV Moves Onto Two DGX Sparks
by Chris DePuy / August 28, 2026 A day after Z.ai opened weights on GLM-5.3-Flash, the desk-side community pushed the model past what I thought was its practical limit on small hardware. Two bottlenecks — the decode path and the KV cache — were attacked in parallel, and both moved in the same 24 hours.…
-
DeepSeek V4 Flash 0731 on the Multi-Node DGX Spark: From One Pair to an Eight-Node Cluster
by Chris DePuy / August 4, 2026 When DeepSeek shipped the official DeepSeek-V4-Flash-0731 weights on July 31, the community benchmarks that followed concentrated on single and dual DGX Spark units. Over the past several days a more interesting pattern has emerged as the same checkpoint moves up the cluster ladder, from a single GB10 to…
-
Poolside’s Laguna S 2.1 118B Punches Above Its Weight
by Chris DePuy / July 28, 2026 This week, the local AI community has been running Poolside’s latest open-weight model, Laguna S 2.1, on everything from single DGX Sparks to 3090 Quad workstations. Released on July 21, Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, a 1M-token…
-

Kimi K3: 2.8 Trillion Parameters, Open Weights — The New Frontier Ceiling
AI-generated concept image of a massive trillion-parameter AI supercomputer. Moonshot AI released Kimi K3 on July 16 — the largest open-weight model ever published at 2.8 trillion total parameters (Moonshot AI announcement). This isn’t a deskside model. It’s the new ceiling the quantization community needs to tackle. Architecture at a Glance Spec Value Total parameters…
-
DGX Spark Clusters by the Numbers: A Sizing Guide Across 1, 2, 3, and 4 Nodes
NVIDIA’s DGX Spark (GB10) started as a deskside curiosity — a 128GB unified-memory workstation drawing ~38W from the GPU during inference. Over the last several weeks, a wave of open-source recipes and community benchmarks has turned the Spark into a modular building block. Users are connecting 2, 3, and 4 units directly — no switch…
