Tag: GGUF
-

1T Inkling Model at 1-bit is 86% smaller
AI-generated concept image of model compression via 1-bit quantization. Thinking Machines Lab’s Inkling — a 975B-parameter (41B active) multimodal with 1M context, vision, and audio (HF model card) — launched day one with a dynamic 1-bit GGUF courtesy of Unsloth. The Compression Format Size Vs. Full BF16 (full) 1.9 TB — NVFP4 ~490 GB 74%…
