Tag: NVFP4
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
-
DFlash2 Beats DSpark for Qwen3.8-27B on One DGX Spark — and Jumps to Apple Silicon
by Chris DePuy / August 20, 2026 A week ago the Qwen3.8-27B decode race on a single NVIDIA DGX Spark was a contest between DSpark and MTP. This week the leading recipes on the same hardware have switched to a different engine: DFlash2, the block-diffusion speculative decoder that first shipped with DeepSeek, is now beating…
-
GLM-5.3-Flash Day Two: DFlash 2 Attacks the Decode Path While 1M-Context KV Moves Onto Two DGX Sparks
by Chris DePuy / August 28, 2026 A day after Z.ai opened weights on GLM-5.3-Flash, the desk-side community pushed the model past what I thought was its practical limit on small hardware. Two bottlenecks — the decode path and the KV cache — were attacked in parallel, and both moved in the same 24 hours.…
-
DeepSeek V4 Flash 0731 on the Multi-Node DGX Spark: From One Pair to an Eight-Node Cluster
by Chris DePuy / August 4, 2026 When DeepSeek shipped the official DeepSeek-V4-Flash-0731 weights on July 31, the community benchmarks that followed concentrated on single and dual DGX Spark units. Over the past several days a more interesting pattern has emerged as the same checkpoint moves up the cluster ladder, from a single GB10 to…
-
Poolside’s Laguna S 2.1 118B Punches Above Its Weight
by Chris DePuy / July 28, 2026 This week, the local AI community has been running Poolside’s latest open-weight model, Laguna S 2.1, on everything from single DGX Sparks to 3090 Quad workstations. Released on July 21, Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, a 1M-token…
-
DeepSeek V4 Flash DSpark on 2× DGX Spark: ~60-67 tok/s with Speculative Decoding
For local inferencing, here is a setup that has proven to be quite stable and fast. It’s the 2x DGX Spark running DeepSeek V4 Flash. It keeps the model running locally and, I’d say, is near Frontier, and doesn’t get lost often when using agents. It’s pretty good at planning, too. The reviews quote ~40…
-

What You Can Run on a DGX Spark Today (Mid-2026)
A wave of open-source recipes and community benchmarks this week clarified what the deskside AI inference market can do in mid-2026 on NVIDIA DGX Spark hardware.
