Recent Blog Posts
-
Kimi K3 Open at 2.8 Trillion Parameters Closes The Gap
by Chris DePuy / July 28, 2026 On July 27, Moonshot AI made good on its promise: the full weights of Kimi K3 — at 2.8 trillion total parameters the largest open-weight model ever released — went live on Hugging Face under a Modified MIT license. The release includes the model’s technical report and a…
-
Poolside’s Laguna S 2.1 118B Punches Above Its Weight
by Chris DePuy / July 28, 2026 This week, the local AI community has been running Poolside’s latest open-weight model, Laguna S 2.1, on everything from single DGX Sparks to 3090 Quad workstations. Released on July 21, Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, a 1M-token…
-

1T Inkling Model at 1-bit is 86% smaller
AI-generated concept image of model compression via 1-bit quantization. Thinking Machines Lab’s Inkling — a 975B-parameter (41B active) multimodal with 1M context, vision, and audio (HF model card) — launched day one with a dynamic 1-bit GGUF courtesy of Unsloth. The Compression Format Size Vs. Full BF16 (full) 1.9 TB — NVFP4 ~490 GB 74%…
-

Kimi K3: 2.8 Trillion Parameters, Open Weights — The New Frontier Ceiling
AI-generated concept image of a massive trillion-parameter AI supercomputer. Moonshot AI released Kimi K3 on July 16 — the largest open-weight model ever published at 2.8 trillion total parameters (Moonshot AI announcement). This isn’t a deskside model. It’s the new ceiling the quantization community needs to tackle. Architecture at a Glance Spec Value Total parameters…
-

GLM-5.2 on 2× DGX Spark: First Non-DeepSeek DSpark Port Hits 19.8 tok/s
AI-generated concept image of dual-processor inference acceleration. Red Hat AI did something the community wasn’t sure was possible: they adapted DSpark speculative decoding to a non-DeepSeek model. GLM-5.2-FP8 now runs with a DSpark draft model, and the numbers are real. The full validation and model weights are published on HuggingFace (RedHatAI/GLM-5.2-speculator.dspark). The Breakthrough DSpark was…
-
DeepSeek V4 Flash DSpark on 2× DGX Spark: ~60-67 tok/s with Speculative Decoding
For local inferencing, here is a setup that has proven to be quite stable and fast. It’s the 2x DGX Spark running DeepSeek V4 Flash. It keeps the model running locally and, I’d say, is near Frontier, and doesn’t get lost often when using agents. It’s pretty good at planning, too. The reviews quote ~40…
-
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters
SGLang vs vLLM vs llama.cpp vs Atlas — The Inference Engine Shootout for DGX Spark Clusters The NVIDIA DGX Spark (GB10) has transformed from a curiosity into a genuine building block for local AI infrastructure. What began as a 128GB unified-memory desktop appliance has, through community effort and NVIDIA’s own software maturation, become a node…
-

Cloudflare is Helping to Paywall the Internet
Cloudflare is Helping to Paywall the Internet The 1990s Internet made information free. AI agents are making it billable. Since the mid-1990s, when the commercial Internet took off, the dominant bargain was simple: content for attention. You browsed, you clicked, you saw ads — and the publisher got paid by advertisers. Most information was free…
-

From Prompts to Loops: Seismic Shift in AI Agent Workflows
A new paradigm shift from manually prompting AI agents to designing autonomous systems that prompt them occurred in June 2026.
-
DGX Spark Clusters by the Numbers: A Sizing Guide Across 1, 2, 3, and 4 Nodes
NVIDIA’s DGX Spark (GB10) started as a deskside curiosity — a 128GB unified-memory workstation drawing ~38W from the GPU during inference. Over the last several weeks, a wave of open-source recipes and community benchmarks has turned the Spark into a modular building block. Users are connecting 2, 3, and 4 units directly — no switch…
Top 10 Companies
2x DGX Spark 802.11ac AI Agents Ajit Pai capex capital spending consolidation DeepSeek Delta Attention DFlash DGX Spark FCC GB10 GLM-5.2 Hugging Face Kimi K3 Local-Inference MoE Moonshot AI MXFP4 Nemotron NVFP4 NVIDIA Open Weights Qwen Speculative Decoding telecommunications Trump Unsloth vLLM
