Tag: DeepSeek
-
NVIDIA’s Official NVFP4 Checkpoints for DeepSeek-V4.1-Flash and GLM-5.3-Flash: Four DGX Sparks Read the Receipts
by Chris DePuy / September 22, 2026 NVIDIA: what the conversion actually changes The DeepSeek checkpoint converts only the ordinary routed experts — projections w1, w2 and w3 across 384 experts in 40 layers — from the source MXFP4 format to NVFP4 weights and activations at group size 16, per the model card; attention, shared…
-
DeepSeek-V4.1-Flash on Four DGX Sparks: The Profile Day One Only Configured Now Boots, Serves and Passes a Million-Token Needle Test
by Chris DePuy / September 14, 2026 Late Sunday evening Pacific time, Mia AI Lab pushed a four-rank serving profile for DeepSeek-V4.1-Flash to its DGX Spark recipe, the same lab whose September 13 post on X put the model on four DGX Sparks and invited readers to run the stack themselves. The significance is not…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
-
DeepSeek V4 Flash 0731 on the Multi-Node DGX Spark: From One Pair to an Eight-Node Cluster
by Chris DePuy / August 4, 2026 When DeepSeek shipped the official DeepSeek-V4-Flash-0731 weights on July 31, the community benchmarks that followed concentrated on single and dual DGX Spark units. Over the past several days a more interesting pattern has emerged as the same checkpoint moves up the cluster ladder, from a single GB10 to…
-
DeepSeek V4 Flash DSpark on 2× DGX Spark: ~60-67 tok/s with Speculative Decoding
For local inferencing, here is a setup that has proven to be quite stable and fast. It’s the 2x DGX Spark running DeepSeek V4 Flash. It keeps the model running locally and, I’d say, is near Frontier, and doesn’t get lost often when using agents. It’s pretty good at planning, too. The reviews quote ~40…
-
DGX Spark Clusters by the Numbers: A Sizing Guide Across 1, 2, 3, and 4 Nodes
NVIDIA’s DGX Spark (GB10) started as a deskside curiosity — a 128GB unified-memory workstation drawing ~38W from the GPU during inference. Over the last several weeks, a wave of open-source recipes and community benchmarks has turned the Spark into a modular building block. Users are connecting 2, 3, and 4 units directly — no switch…
-

What You Can Run on a DGX Spark Today (Mid-2026)
A wave of open-source recipes and community benchmarks this week clarified what the deskside AI inference market can do in mid-2026 on NVIDIA DGX Spark hardware.
