Recent Blog Posts
-
New TensorFold LLM Model Serving Software Dramatically Improves Decode Speed on Macs and Sparks
by Chris DePuy, 650 Group / September 27, 2026 Two repositories shipped from one maintainer this week and made the same promise in two different places. On September 25 and 26, a developer publishing under the name Ash Hart released TensorFold, a local inference engine for Apple Silicon and NVIDIA GPUs, and Imprint, a tool…
-
NVIDIA’s Official NVFP4 Checkpoints for DeepSeek-V4.1-Flash and GLM-5.3-Flash: Four DGX Sparks Read the Receipts
by Chris DePuy / September 22, 2026 NVIDIA: what the conversion actually changes The DeepSeek checkpoint converts only the ordinary routed experts — projections w1, w2 and w3 across 384 experts in 40 layers — from the source MXFP4 format to NVFP4 weights and activations at group size 16, per the model card; attention, shared…
-
A Cross-Silicon Inference Loop: Two DGX Sparks, Two Mac Studios, and an Open RDMA Bridge
by Chris DePuy / September 19, 2026 On September 17, the community project pd-bridge published the result its earlier one-way design had been pointed at: a working closed loop in which two NVIDIA DGX Sparks prefill a 284-billion-parameter DeepSeek model, two Apple Mac Studios receive and serve the resulting attention state, and a multi-turn conversation…
-
The System One Split: Jev, CUA-S1-FORMS, and the Decision Layer Coming Beneath the Agent Loop
by Chris DePuy / September 20, 2026 Between September 15 and 18, four related releases arrived within four days. TypeSafe AI emerged from two years of stealth with Jev, a model class the company calls System One, LangChain published an integration guide showing Jev wired into agent middleware, Browser Use open-sourced Jev Ultrafast, a browser-agent…
-
DeepSeek-V4.1-Flash on Four DGX Sparks: The Profile Day One Only Configured Now Boots, Serves and Passes a Million-Token Needle Test
by Chris DePuy / September 14, 2026 Late Sunday evening Pacific time, Mia AI Lab pushed a four-rank serving profile for DeepSeek-V4.1-Flash to its DGX Spark recipe, the same lab whose September 13 post on X put the model on four DGX Sparks and invited readers to run the stack themselves. The significance is not…
-
Cloudflare’s New AI Crawler Defaults Are a Concrete Step Toward 402
Cloudflare is changing the way we get information. It is doing it quietly, one settings migration at a time. On September 15, Cloudflare began changing the recommended setting for AI training crawlers from Block to “Disallow AI Training” — a distinction that lets major search crawlers such as Applebot, Googlebot and Bingbot continue indexing a…
-
GLM-5.3-Flash on a Single DGX Spark: Three Community Quantizations Fit Where the Vendor’s Does Not
by Chris DePuy / September 13, 2026 On September 9, NVIDIA published an official NVFP4 quantization of Z.ai’s GLM-5.3-Flash. Its Model Optimizer recipe leaves the key-value cache at FP8 and compresses only the linear layers inside the transformer blocks — the sparse-MoE shared experts and the dense MLP — from 16 bits down to 4,…
-
Day One on Three and Four Sparks: Community Recipes Put DeepSeek-V4.1-Flash on the Desk
by Chris DePuy / September 12, 2026 On September 10, DeepSeek released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with a 552-billion-parameter backbone and a context window of one million tokens. The launch kits for NVIDIA’s desk-side DGX Spark machines arrived the same day: a four-machine stack on vLLM was serving requests by 9:17 Eastern time, roughly…
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
Top 10 Companies
AI Agents AI Infrastructure DeepSeek DeepSeek V4.1-Flash Deskside AI DFlash DFlash2 DGX Spark DSpark FP8 GB10 GLM-5.3-Flash inference benchmark Kimi K3 KV Cache Local-Inference long context MoE MXFP4 Nemotron NVFP4 NVIDIA Open Weights Quantization Qwen SGLang Speculative Decoding Tensor Parallelism Unified Memory vLLM
