New TensorFold LLM Model Serving Software Dramatically Improves Decode Speed on Macs and Sparks

Written by

in

by Chris DePuy, 650 Group / September 27, 2026

Two repositories shipped from one maintainer this week and made the same promise in two different places. On September 25 and 26, a developer publishing under the name Ash Hart released TensorFold, a local inference engine for Apple Silicon and NVIDIA GPUs, and Imprint, a tool that saves and restores the model state behind a repeated context. Both are pitched as speed, and both attach a stronger claim underneath the speed: that the fast path returns exactly what the slow path would have returned. In the same week, the Spark serving kits whose throughput panels we have been tracking began publishing quality gates alongside their numbers. I think that pairing, more than any single token rate, is the development worth recording. In our own tests of TensorFold, on a Mac, prefill does not improve (compared to using LM Studio), but decode is like 3 times faster. I’m really impressed with Ash’s work.

Release cadence is the first signal. TensorFold’s tagged releases run from 0.2.0 on September 25 at 22:45 UTC to 0.3.4.1 on September 27 at 08:47 UTC — seven versions in roughly thirty-four hours, the last of them a fix for a regression the project introduced the same evening. The code is MIT-licensed, installs with pip, and serves /v1/chat/completions on the standard OpenAI surface, reading its weights directly from Hugging Face.

Two different paths of light through glass arriving at one identical point on a single dark plate
Two routes, one result. Concept illustration, generated locally with ComfyUI (SDXL). Source: generated for Market Intelligence Research.

TensorFold: the contract is what it does not change

The engine’s own statement of its claim is specific and modest: on DGX Spark the project measures 1.6 to 3x the decode rate of vLLM when both sides run multi-token-prediction drafts, on one Spark or two. What makes that number usable rather than promotional is the pair of rules the README states in support of it. First, the token at each position is the argmax over the request’s top-k and top-p candidates of the row’s logit divided by temperature, plus a Gumbel term obtained by hashing together the seed, the position in the sequence and the token id — an exact draw from the distribution those logits imply, dependent on nothing else. Second, a row inside a verification pass over several drafted tokens receives the identical bits it would receive as a single-row step. Together they mean a drafted token survives exactly when it equals what serial decoding would have sampled, so the drafted path writes the same bytes the serial path would have, and the drafts change speed only. The project’s own self-check is a request sent with "draft": false to decode one token a round for comparison.

The published comparison is a single stream of 64-token replies through the same OpenAI client against vLLM with MTP=3, code prompts, sampled and greedy, with every TensorFold figure holding byte-for-byte equality against the project’s own serial path.

Configuration (one stream, code)TensorFoldvLLM MTP=3Ratio
Qwen3.8-27B + DFlash2, 1 Spark49.617.72.8x
Qwen3.8-27B + DFlash2, 2 Sparks82.433.12.5x
Qwen3.8 Flash Next, 1 Spark68.342.41.6x
Qwen3.8 Flash Next, 2 Sparks103.846.42.2x
GLM-5.3-Flash, 2 Sparks49.424.52.0x
GLM-5.3-Flash, 2 Sparks, greedy66.332.22.1x

All figures are decode tokens per second measured by the project on DGX Spark through the same client. Apple Silicon carries the same lane engine across the line. On an M5 Max with 128 GB, the 0.3.4 benchmark put Nemotron 3.5 Lightning at 288, 223, 292 and 243 tok/s across four 64-token cells against 173, 173, 179 and 176 for MLX-LM’s own server through the same client, and Qwen3.8-27B with the DFlash2 drafter at 120 to 124 tok/s on a short answer with thinking and 189 on code, against 27 and 26 without drafts. Lane batching now runs on every Apple Silicon generation; on M1 through M4, which lack the Metal tensor units, the project substitutes its own row-exact lane decoder plus simdgroup matmul kernels and measures 1.9x to 4x the serial rate on an M3 Ultra. One licensing constraint is worth flagging for anyone deploying this commercially: the GLM-5.3-Flash lane’s optional draft model, incoai’s DFlash2, is licensed for non-commercial use only under CC BY-NC-ND 4.0, and without it that model drafts from its own head alone.

The regression it published against itself

The more informative document this week is not the speed table but the 0.3.4.1 release note. Version 0.3.4 had routed the prompts of Nemotron and Flash Next through its row-exact decode kernels, so a resumed conversation would reproduce a fresh conversation bit for bit — and the price was prefill running several multiples slower than MLX’s: roughly 550 tok/s for Nemotron on an M5 Max, where MLX-LM runs 2,500 to 3,000, and roughly 300 on an M3 Ultra for Flash Next, against 760 for MLX. The fix returns every model’s prefill to MLX’s own forward, processed in chunks aligned to a fixed grid of 2,048 tokens, with prompt caches kept only at grid points so that a chunk’s bits are the same regardless of where a conversation resumed. Cold-prompt processing then measured 3,400 tok/s at 8k for Nemotron on the M5 Max against 2,540 for MLX-LM in the same session, 2,456 at 32k against 2,448, and 934 at 2k for Flash Next on the M3 Ultra.

The equivalence checks in that note are the part worth holding onto. Every drafted reply equaled the same request sent with drafting disabled, nine of nine on each model; resumed multi-turn conversations matched the same exchanges started from scratch, among them a three-turn Flash Next chat picked up at nine grid points inside a 19k-token opening message; and the suite passes 370 tests. One limit is stated rather than buried: on Flash Next and Nemotron, which prefix happened to sit in the cache can shift a prompt’s final bits, and with them the reply — though the drafted and serial paths agree whenever they start from the same cache.

Imprint: the same question asked about the wait before the first token

Imprint attacks the other half of the delay. It computes a repeated context once, saves the resulting model state, and restores it after the model shuts down, so a matching request processes only its new suffix — the shape of an agent harness that resends its instructions and project files every session. The headline is a reported time-to-first-token of 19 seconds falling to 0.3 seconds, about 63x. The project’s own evidence file then does something rarer than the claim: it states that the figure is author-reported, that raw timing traces, workload settings and repeat counts are not included, that it has not yet been reproduced independently using the standalone CLI, and that the package’s model-free tests “do not prove numerical equivalence, real model memory release or end-to-end latency.” It also declines to claim novelty, describing the work as practical engineering around existing prompt-cache concepts rather than a new theory of KV-cache portability. The repository also carries no licence file, which for a commercial operator is the first question to settle rather than the last.

One lab has since put its own numbers on a comparable path. The Volatile Markets roundup, written from a Mac Studio and DGX Spark lab that built a Mac Studio backend for the tool, reports a 44K-token context whose cold processing took 188 seconds coming back in under two seconds and bit-exact, and reports the same lab’s TensorFold build taking GLM-5.3-Flash on an M3 Ultra from 45 to 60 tok/s with word-for-word identical answers. Both of those are that lab’s measurements rather than the project’s, and they are the kind of independent replication the evidence file is waiting for.

The Spark kits added their own gates

The same discipline appeared across the Spark serving recipes this week, in three places. On a single GB10, the Qwen3.8-Flash-Next recipe moved to vLLM 0.30 with five sampled draft tokens checked by an exact probability-ratio test, an 830,582-token bf16 KV pool and sixteen seats: 82 tok/s peak at one stream against 68 for the previous version, and 318 at sixteen streams, booting in about four minutes instead of twelve. The maintainer’s position on quality is explicit — output cannot change, because the generated text follows the model’s own distribution under either setting — and the default recipe pairs an opt-in bit-exact decoding mode for evaluations and debugging with a faster lane that reads NVIDIA’s official weights at the cost of about 25% slower prose.

On four Sparks, the DeepSeek-V4.1-Flash TP4 profile reports a fresh-clone fleet booting healthy in 160 to 180 seconds, a 5.98-to-6.49-million-token KV pool spanning a 1M-token context, prose decode of 89.76 tok/s at one stream, and about 4,800 tok/s of real-text prefill from a 15k prompt out to 123k — with needle retrieval passing at 1,011,084 tokens. Its correctness arguments are equally specific: the greedy transcript hash-checked byte-identical to the prior version’s across fifteen outputs at the first boot and six at the second, every draft token checked against the target model, and each adapter either reproducing the stock path’s bits exactly or matching it in distribution. The maintainer also publishes where that stops being true, noting that the new all-reduce sums its partials differently from NCCL, so the arithmetic is not bit-identical across the two transports — which is why the maintainers score the profile as well.

The GLM-5.3-Flash TP4 profile shows what that scoring looks like when it does not fully pass. Prose decode runs 71.64, 149.44 and 305.59 tok/s at one, four and sixteen concurrent requests, with about 2,200 tok/s of fresh prefill between 16k and 64k prompts — and against an accepted quality bar of 75/75, the profile scores qeval 72/75 at one stream and a teacher-forced mean KL of 0.029189 against an accepted 0.028837 over seventeen items and 6,618 positions. The maintainers conclude that the profile clears its predefined validation gates without establishing that quality is unchanged. A separate 4.75-bit EXL3 build of the DeepSeek model reaches 64.8 tok/s of code decode at one stream and 166.7 aggregate across ten, and the two kernel changes behind the gain are each described as output bit-identical.

What it leaves the operator

I think the equivalence statement is what turns a throughput number into something an operator can act on. A speedup that changes the output is a new model to re-evaluate; a speedup that provably writes the same bytes is a configuration change, and it can be regression-tested against a stored baseline the way any other infrastructure change is. That is also where the gates earn their place: a kit that publishes an acceptance bar and then reports landing under it, as the four-Spark GLM profile does, is more useful than one that publishes only the fast column, and a tool that lists what its tests do not prove is more useful than one that lists only what they do. We treat single-maintainer measurements as directional until an independent replication lands, and in this layer the replications are arriving faster than they did a month ago. As firms such as 650 Group have been tracking in their AI infrastructure research, the cost and control points of the local serving stack keep migrating away from the model and into the machinery around it, and this week that machinery began certifying its own output.

More posts

© MarketIntelligenceResearch.com