1T Inkling Model at 1-bit is 86% smaller

Large cube compressed to small cube representing 1-bit quantization

Written by

in

AI-generated concept image of model compression via 1-bit quantization.
AI-generated concept image of model compression via 1-bit quantization.

Thinking Machines Lab’s Inkling — a 975B-parameter (41B active) multimodal with 1M context, vision, and audio (HF model card) — launched day one with a dynamic 1-bit GGUF courtesy of Unsloth.

The Compression

Format Size Vs. Full
BF16 (full) 1.9 TB
NVFP4 ~490 GB 74% smaller
UD-IQ1_M (1-bit, recommended) 270 GB 86% smaller

From 1.9 TB to 270 GB with 74.2% top-1% accuracy retention (Unsloth docs). That’s a model that originally required cluster-level storage now fitting on a Mac Studio Ultra (~290 GB RAM threshold) or a multi-GPU workstation. Both UD-IQ1_S and UD-IQ1_M variants are available on HuggingFace; Unsloth recommends UD-IQ1_M for the best balance of accessibility and accuracy.

What Survives

Inkling accepts text, images, and audio — all three modalities survive the 1-bit quantization (Thinking Machines blog). The model excels at:

  • Coding and agentic/tool-calling
  • RAG and chat
  • Multilingual applications
  • Multimodal reasoning (text + vision + audio)

The Full Quantization Ladder

Quant Size Target Hardware
UD-IQ1_M (rec.) 270 GB Mac Studio Ultra, 2× RTX 6000
UD-IQ1_S ~255 GB Max memory savings (slightly lower accuracy)
UD-IQ2_XXS ~310 GB High-end workstation
UD-IQ3_XXS ~350 GB Workstation
Q4_K_M ~520 GB RTX 6000-class
Q6_K ~750 GB Multi-GPU
Q8_0 ~950 GB Cluster
NVFP4 ~490 GB Near-lossless path

All sizes confirmed from Unsloth’s GGUF config.

Quick Takeaways

Unsloth’s Dynamic 2.0 quantization is weight-level — it preserves the most important parameters at higher precision while aggressively quantizing the rest. The result: an 86% size reduction that keeps Inkling genuinely useful for coding, multimodal reasoning, and agentic tasks.

Available now on HuggingFace: unsloth/inkling-GGUF. Apache 2.0 licensed (HF model card). Also available on Baseten’s inference platform for managed deployment.


Sources: Thinking Machines Lab on HuggingFace, Unsloth Inkling GGUF quantized models, Unsloth documentation, Thinking Machines Lab announcement.

More posts

© MarketIntelligenceResearch.com