Tag: Agentic AI
-
GLM-5.3-Flash on a Single DGX Spark: Three Community Quantizations Fit Where the Vendor’s Does Not
by Chris DePuy / September 13, 2026 On September 9, NVIDIA published an official NVFP4 quantization of Z.ai’s GLM-5.3-Flash. Its Model Optimizer recipe leaves the key-value cache at FP8 and compresses only the linear layers inside the transformer blocks — the sparse-MoE shared experts and the dense MLP — from 16 bits down to 4,…
-
Nemotron 3.5 Lightning: NVIDIA Carves Out an Agent Execution Layer for the Desk-Side
by Chris DePuy / August 12, 2026 NVIDIA this week released Nemotron 3.5 Lightning, an open 30B MoE model with just 3B active parameters, and I think it is the clearest signal yet that the deskside inference market is splitting into two distinct tiers. The model is built for the high-volume execution layer of long-running…
-
Grok Bot, Two Weeks In: The Always-On Agent Workforce Gets a Shared Computer and a Shared Security Boundary
by Chris DePuy / August 29, 2026 xAI launched Grok Bot in early beta on August 11, and in the two weeks since, the community has made the platform’s ambition concrete faster than the company’s own launch materials do. The product is not another chat interface — a Bot gets its own cloud computer, signs…
