Inkling
Release in the Inkling family · version Inkling
Maintained by Thinking Machines Lab13
Inkling is a 66-layer decoder-only mixture-of-experts transformer with 975B total and 41B active parameters, routing each token to 6 of 256 experts plus 2 shared experts. It accepts text, image, and audio input, outputs text, and supports a context window of up to 1M tokens. Thinking Machines Lab says it trained the model from scratch on 45 trillion tokens of text, images, audio, and video.12
- Model hub: Hugging Face (BF16) (external site: huggingface.co)
- Model hub: Hugging Face (NVFP4) (external site: huggingface.co)
- Documentation: Inkling Model Card (external site: thinkingmachines.ai)
- Release notes: Inkling: Our Open-Weights Model (external site: thinkingmachines.ai)
- License: Model Acceptable Use Policy (external site: thinkingmachines.ai)
Availability and license
Overall availability
BF16 and NVFP4 checkpoints are downloadable from Hugging Face without gating; the developer's Model Acceptable Use Policy also applies.34
Availability is separate from permission: read the license before using or redistributing.
The Hugging Face repository declares Apache 2.0 in its metadata and links to the Apache license text; it has no separate LICENSE file. Thinking Machines Lab also publishes a Model Acceptable Use Policy (July 15, 2026) stating that anyone who accesses, downloads, or uses its model materials agrees to be bound by it. The policy lists prohibited uses such as weapons development, CSAM, fraud, and unauthorized surveillance, and does not mention the Apache license.34
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published as a BF16 checkpoint and a separate NVFP4-quantized checkpoint.31 |
| Inference codeIs code for running the model published? | Public | The model card documents deployment with open-source SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face Transformers, with published recipes.13 |
| Training codeIs the code used to train the model published? | Unknown | The Tinker Cookbook is published for fine-tuning; no code for training Inkling itself was found. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card and training-data documentation describe source categories (public web content and repositories, data acquired from third parties or partners, and synthetic data), deduplication, and filtering. The data is not released.153 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The announcement describes the architecture, the optimizer setup (Muon for large matrices, Adam for other parameters), weight-decay scheduling, an SFT bootstrap, and large-scale reinforcement learning at a high level.2 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | Benchmark and safety results are published in the model card and as evaluation result files in the Hugging Face repository. The announcement notes that some reported results came from a different checkpoint than the one released.32 |
What it is useful for
Run and use notes
- The model card states that the BF16 checkpoint needs at least 2 TB of aggregated GPU memory (for example 8x NVIDIA B300 or 16x H200), and that the NVFP4-quantized checkpoint needs at least 600 GB (W4A4 on 4x B300, which requires SM100+ architecture, or W4A16 on 8x H200).1
- Documented deployment frameworks are SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face Transformers; the model is also available through Tinker and third-party inference providers.13
Organization context
Provenance and derivatives
Thinking Machines Lab says it trained Inkling from scratch. Its announcement states that post-training began with supervised fine-tuning on synthetic data generated by open-weights models including Kimi K2.5, which it says was a small fraction of compute, followed by large-scale reinforcement learning. No other model's weights are named as a starting point.2
Catalog records that name this entry in their provenance:
Other releases in the Inkling family
- Inkling-SmallModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.