Introducing Inkling (Thinking Machines Lab)
Thinking Machines Lab‘s first-party announcement of Inkling
(2026-07-15), read alongside its model card (thinkingmachines.ai/model-card/inkling/) and pulled in
to ground the Willison writeup that routed here.
What it supplies
The hard numbers the secondary coverage doesn’t carry: the 66-layer decoder-only MoE architecture
(256 routed + 2 shared experts, 6 active per token; sigmoid router with auxiliary-loss-free balancing;
5:1 sliding-window-to-global attention; relative positional embeddings instead of RoPE; short convolutions
in the attention and residual paths), the 1M-token context, the encoder-free vision (40×40 patches,
hMLP) and audio (dMel spectrograms) paths, the 30M+ RL rollouts in post-training, the full
benchmark table at effort=0.99, the BF16 2TB+ / NVFP4 ~600GB VRAM figures, and the serving and
fine-tuning surfaces. All of it is on inkling.
Positioning, in their words
“A broad, balanced foundation model: strong across many domains, flexible enough to adapt” — and, plainly, “Inkling is not the strongest overall model available today, open or closed.” The differentiators claimed are multimodality, efficient thinking, and Tinker fine-tuning, aimed at organizations building custom models rather than buyers shopping a leaderboard.
Tier & provenance
T1 — the maker’s own announcement and model card. First-party, which means the benchmarks are
self-reported and the positioning is marketing: authoritative on what the model is, not on how good it
is. No independent artificial-analysis re-run existed at ingest. The training-data documentation page
it links stays vague (“publicly accessible data repositories”), a gap Willison
also flags. freshness: volatile — pricing carries a limited-time 50% discount and the benchmark table
will age.
Related
inkling · thinking-machines-lab · tinker · inkling-willison-review · moe-architecture · open-weight-models · llm-benchmarks