Tag: MXFP4
-
NVIDIA’s Official NVFP4 Checkpoints for DeepSeek-V4.1-Flash and GLM-5.3-Flash: Four DGX Sparks Read the Receipts
by Chris DePuy / September 22, 2026 NVIDIA: what the conversion actually changes The DeepSeek checkpoint converts only the ordinary routed experts — projections w1, w2 and w3 across 384 experts in 40 layers — from the source MXFP4 format to NVFP4 weights and activations at group size 16, per the model card; attention, shared…
-
DeepSeek-V4.1-Flash on Four DGX Sparks: The Profile Day One Only Configured Now Boots, Serves and Passes a Million-Token Needle Test
by Chris DePuy / September 14, 2026 Late Sunday evening Pacific time, Mia AI Lab pushed a four-rank serving profile for DeepSeek-V4.1-Flash to its DGX Spark recipe, the same lab whose September 13 post on X put the model on four DGX Sparks and invited readers to run the stack themselves. The significance is not…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
-
Kimi K3 Open at 2.8 Trillion Parameters Closes The Gap
by Chris DePuy / July 28, 2026 On July 27, Moonshot AI made good on its promise: the full weights of Kimi K3 — at 2.8 trillion total parameters the largest open-weight model ever released — went live on Hugging Face under a Modified MIT license. The release includes the model’s technical report and a…
-

What You Can Run on a DGX Spark Today (Mid-2026)
A wave of open-source recipes and community benchmarks this week clarified what the deskside AI inference market can do in mid-2026 on NVIDIA DGX Spark hardware.
