Tag: NVIDIA
-

GLM-5.2 on 2× DGX Spark: First Non-DeepSeek DSpark Port Hits 19.8 tok/s
AI-generated concept image of dual-processor inference acceleration. Red Hat AI did something the community wasn’t sure was possible: they adapted DSpark speculative decoding to a non-DeepSeek model. GLM-5.2-FP8 now runs with a DSpark draft model, and the numbers are real. The full validation and model weights are published on HuggingFace (RedHatAI/GLM-5.2-speculator.dspark). The Breakthrough DSpark was…
-
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters The NVIDIA DGX Spark (GB10) has transformed from a curiosity into a genuine building block for local AI infrastructure. What began as a 128GB unified-memory desktop appliance has, through community effort and NVIDIA’s own software maturation, become a node…
-
DGX Spark Clusters by the Numbers: A Sizing Guide Across 1, 2, 3, and 4 Nodes
NVIDIA’s DGX Spark (GB10) started as a deskside curiosity — a 128GB unified-memory workstation drawing ~38W from the GPU during inference. Over the last several weeks, a wave of open-source recipes and community benchmarks has turned the Spark into a modular building block. Users are connecting 2, 3, and 4 units directly — no switch…
-

What You Can Run on a DGX Spark Today (Mid-2026)
A wave of open-source recipes and community benchmarks this week clarified what the deskside AI inference market can do in mid-2026 on NVIDIA DGX Spark hardware.
