Tag: benchmark
-
Qwen3.8-Flash-Next on One DGX Spark: NVIDIA’s Official NVFP4 Weights Hit 44-49 tok/s, and Scaling Buys Context, Not Speed
by Chris DePuy / September 7, 2026 Two weeks ago, Qwen3.8-Flash-Next was a day-zero curiosity: NVIDIA’s official NVFP4 quantization of the model landed on Hugging Face on August 31, and the community was still proving it could run at all. This past week it stopped being a curiosity. Three independent reference deployments now serve the…
