by Chris DePuy, 650 Group / October 6, 2026
Five DGX Sparks now carry the full 753-billion-parameter GLM-5.3: an operator reported the first working five-board tensor-parallel build on Monday, with 1.5 million tokens of KV cache, prefill above 2,000 tokens per second, and 76 tokens per second single-stream — rising to 98–113 tokens per second during long reasoning passes and 272 aggregate at eight concurrent requests, per the operator’s own build note. The same day, Reflection AI introduced Beam, a 501-billion-parameter open-weight model whose selling point is not a benchmark ceiling but inference economics: scores comparable to GLM-5.2 while using an estimated 3–4× less inference compute, per the company’s announcement. I think the two arrivals frame the same question from opposite ends of the budget: once a five-board desk cluster can serve a frontier checkpoint at hundreds of tokens per second, the race moves to intelligence per token — and both the model labs and the Spark cluster operators are now publishing on that axis.
The five-board build: full GLM-5.3 at 1.5 million tokens of cache
The five-Spark lane passed through a quality screen before it shipped a throughput figure: the operator reports 10 of 10 on an Estonia-style verification set and 9 of 10 on a long-context reasoning probe (five exact, four near, one miss) — informal checks, self-reported, not the audited screens the maintainer recipes publish, but the first numbers attached to a TP5 build. The decode ladder runs 76 tokens per second single-stream, 98–113 on long reasoning, and 272 aggregate at eight concurrent requests, with prefill above 2,000 tokens per second, and the operator marks the build “not optimized.” Five boards is a new data point on the cluster-size curve: two Sparks serve GLM-5.3-Flash at 57–60 tokens per second single stream, four serve the full model, and now five serve it with a KV budget of 1.5 million tokens — context capacity scaling roughly linearly with boards while decode scales on the all-reduce path.
The scaling context comes from the recipes published the same week. The Mia AI Lab two-Spark kit measured single-stream prose at 60.4 tokens per second, 108.8 aggregate at four streams, and 130.8 at eight (167.0 on code) with the full 1,048,576-token window per request over a shared 2.9-million-token FP8 KV pool; its 981,841-token needle prompt read at 1,015 tokens per second with the needle found. A three-Spark appliance rebuilt on a TensorFold engine measured code decode rising from 61.1 to 91.3 tokens per second against its own prior release and cold 64k prefill at 2,124 tokens per second, with conversation state persisted to disk — a 100,000-plus-token chat resumes in 1.5 seconds against 49.9 seconds cold. An independent operator’s two-board replication of the Mia checkpoint landed at 57.4 tokens per second standard and 55.3 abliterated, median single-stream decode — within about ten percent of the maintainer’s own table. The TP5 figure sits where the curve implies it should, and its decode-per-request drops steeply at eight streams (272 aggregate), which is the all-reduce tax the four-board fabric work below is trying to erase.
The fabric layer: a 17-microsecond all-reduce and what is still unproven
The deepest layer of the stack got its own release. A new Apache-2.0 library, dgx-spark-networking, packages switchless RoCE collectives for two to four Sparks wired directly over their ConnectX-7 ports with no switch: cabling four boxes into a ring, teaching NCCL to stay on neighbor edges for the bulk collectives, and shipping a one-shot RDMA all-reduce that finishes a 10-KiB reduction in 17 microseconds against NCCL’s 91 — one kernel, replayable inside CUDA graphs, bit-identical on every rank. The measured TP4 profile balances every interface between 24.8 and 25.2 percent of the fabric’s RDMA bytes with prefill up 1.6 to 1.9 percent and decode within noise.
The library is equally explicit about what has not happened yet: the vendored one-shot runtime has not been run end-to-end through this package on real Sparks, the serving integration still points at the b12x tree, and the GPU-initiated transport — where the GPU would ring the NIC’s doorbell itself — is staged rather than implemented, after a first attempt rebooted a node. A DOCA GPUNetIO path with the host proxy is qualified next; its write-latency sample ran on every link of four Sparks at 4.1 to 7 microseconds half round trip. A 17-microsecond collective aimed at the next-token path targets exactly the aggregate-decode drop the TP5 build exhibits, but until a serving recipe drives it end to end, the receipts are the fabric’s own — the honest state of a layer that is one release ahead of its application.
Beam: the efficiency axis arrives as a product claim
Reflection AI introduced Beam on October 5: a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, built for coding, reasoning, and agentic workloads, with full weights under an Apache 2.0 license promised later this month per the company’s announcement, with quantized FP8 and NVFP4 numerics named in the release thread for deployment. The company’s table shows Beam ahead of Inkling on all four coding tests where both report — 77.2 versus 56.9 on SWE Bench Pro v2-Hard, 80.9 versus 77.6 on SWEBench Verified — and lands at parity with GLM 5.2 on DeepSWE (44.4 versus 44.0), behind GLM 5.3 (61.0) and Kimi K3 (68.0) there. The efficiency claim is the differentiator: scores comparable to GLM-5.2 while using 3–4× less inference compute, estimated from generation forward-pass FLOPs and generated-token counts against third-party benchmark data — a vendor estimate pending replication.
The reinforcement-learning run is the scale story. Reflection deployed 10,500 NVIDIA GB300 GPUs for four weeks, generated more than 100 million rollouts at up to 256K context, ran roughly 1.3 billion sandboxes, sustained an average of 110,000 concurrent rollouts, and drew from a pool of nearly one million environments; the company calls it one of the largest RL runs by any open lab to date. Pretraining ran end-to-end in under four weeks on 6,144 GB300 NVL72 GPUs over 23.8 trillion tokens, finishing at 92.3 percent goodput, and the training method is fully asynchronous policy gradients that stay stable when rollouts arrive more than a day stale — 107 weight versions behind the live policy. For comparison, Inkling was trained on 30 million rollouts three months earlier. I treat RL compute, not parameter count, as the axis where the open labs are currently compounding — and Beam’s 3–4× efficiency framing is the first time a US open-weight lab has made tokens-per-dollar the headline rather than the footnote.
What the two lanes share
The overlap is uncited but real: Beam’s differentiator is inference compute per task, and the Spark cluster tier is the natural place where such a model’s economics get tested by the operators who pay in watts and boards rather than API credits. I read the two arrivals as one market event in two registers: the labs pricing intelligence per token, the operators proving how many tokens a fixed pool of boards and watts can carry. The five-board build shows the desk tier absorbing a 753-billion-parameter frontier checkpoint with a 1.5-million-token cache; a 501-billion-parameter model with 23 billion active would sit comfortably inside the four-board lane’s memory budget, and its FP8 and NVFP4 numerics — formats the desk-side engines added support for in the past two releases — are named in the release thread. The verification points for both lanes are concrete: whether a serving recipe adopts the 17-microsecond collective and what it does to aggregate decode at eight streams, and whether third-party runs reproduce Beam’s efficiency ratio on measured serving stacks. Every cluster figure in this post is a single-maintainer or single-operator measurement unless marked as replication — the standing caveat for this tier — and the Beam benchmark table is the vendor’s own pending independent evaluation. As firms such as 650 Group have been tracking in their AI infrastructure research, the desk-side question has moved from which model fits on a cluster to which cluster serves a frontier model at the lowest cost per useful token — and this week both a five-board Spark build and a 501-billion-parameter efficiency-first model answered on that axis, receipts attached.
