Tag: GB10
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
-
GLM-5.3-Flash Day Two: DFlash 2 Attacks the Decode Path While 1M-Context KV Moves Onto Two DGX Sparks
by Chris DePuy / August 28, 2026 A day after Z.ai opened weights on GLM-5.3-Flash, the desk-side community pushed the model past what I thought was its practical limit on small hardware. Two bottlenecks — the decode path and the KV cache — were attacked in parallel, and both moved in the same 24 hours.…
-
DeepSeek V4 Flash 0731 on the Multi-Node DGX Spark: From One Pair to an Eight-Node Cluster
by Chris DePuy / August 4, 2026 When DeepSeek shipped the official DeepSeek-V4-Flash-0731 weights on July 31, the community benchmarks that followed concentrated on single and dual DGX Spark units. Over the past several days a more interesting pattern has emerged as the same checkpoint moves up the cluster ladder, from a single GB10 to…
-
Poolside’s Laguna S 2.1 118B Punches Above Its Weight
by Chris DePuy / July 28, 2026 This week, the local AI community has been running Poolside’s latest open-weight model, Laguna S 2.1, on everything from single DGX Sparks to 3090 Quad workstations. Released on July 21, Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, a 1M-token…
-
DeepSeek V4 Flash DSpark on 2× DGX Spark: ~60-67 tok/s with Speculative Decoding
For local inferencing, here is a setup that has proven to be quite stable and fast. It’s the 2x DGX Spark running DeepSeek V4 Flash. It keeps the model running locally and, I’d say, is near Frontier, and doesn’t get lost often when using agents. It’s pretty good at planning, too. The reviews quote ~40…
-
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters The NVIDIA DGX Spark (GB10) has transformed from a curiosity into a genuine building block for local AI infrastructure. What began as a 128GB unified-memory desktop appliance has, through community effort and NVIDIA’s own software maturation, become a node…
-
DGX Spark Clusters by the Numbers: A Sizing Guide Across 1, 2, 3, and 4 Nodes
NVIDIA’s DGX Spark (GB10) started as a deskside curiosity — a 128GB unified-memory workstation drawing ~38W from the GPU during inference. Over the last several weeks, a wave of open-source recipes and community benchmarks has turned the Spark into a modular building block. Users are connecting 2, 3, and 4 units directly — no switch…
-

What You Can Run on a DGX Spark Today (Mid-2026)
A wave of open-source recipes and community benchmarks this week clarified what the deskside AI inference market can do in mid-2026 on NVIDIA DGX Spark hardware.
