Cogito v2.1 671B
Release in the Cogito family · version cogito-671b-v2.1
Maintained by Deep Cogito14
A 671B-parameter mixture-of-experts model with 37B active parameters that Deep Cogito post-trained in-house from the DeepSeek-V3 base model. It is a hybrid reasoning model that can answer directly or reason first, supports a 128k-token context, and was trained in over 30 languages; Deep Cogito says process supervision of reasoning chains lets it reach answers with shorter reasoning.145
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Release notes: Cogito v2.1 announcement (external site: deepcogito.com)
- Repository: Cogito v2 and v2.1 repository (external site: github.com)
Availability and license
Overall availability
Weights download from Hugging Face without an access gate; an FP8 version is published separately. The announcement also lists API access through third-party platforms and a free web chat.124
Availability is separate from permission: read the license before using or redistributing.
The model card states that the repository and the model weights are licensed under the MIT License, and the companion GitHub repository carries an MIT LICENSE file; the Hugging Face repository has no separate LICENSE file. Deep Cogito calls its DeepSeek starting point an open-licensed base model; the DeepSeek-V3-Base model card says its code is MIT-licensed and that use of the DeepSeek-V3 Base and Chat models is subject to DeepSeek's Model License.13647
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | BF16 safetensors in an ungated Hugging Face repository; an FP8 repository is also published.12 |
| Inference codeIs code for running the model published? | Public | The repository includes DeepSeek modeling code, and the model card gives Transformers and vLLM examples, including tool calling.13 |
| Training codeIs the code used to train the model published? | Unknown | No training code was found in the model card, announcement, or GitHub repository, which contains usage examples. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Unknown | The model card and announcement do not describe the post-training data; the DeepSeek-V3 base model's pretraining data is described only in DeepSeek's materials. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The announcement describes the approach only at a high level (forking the DeepSeek base model, in-house post-training, process supervision of reasoning chains); the model card names Iterated Distillation and Amplification. No configurations are given.41 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The model card and announcement show benchmark charts and list repeats per example for each benchmark; evaluation code is not linked.14 |
What it is useful for
The model card describes it as optimized for coding, STEM, instruction following, general helpfulness, and tool calling.1
Run and use notes
- The model card states that the BF16 checkpoint takes about 1.3 TB for parameters and needs at least 8 B200 GPUs (one node) or 16 H200 GPUs (two nodes), and points users with 8 H200 GPUs to the FP8 checkpoint. Its vLLM example uses tensor parallel size 8.1
Organization context
Provenance and derivatives
Post-trained in-house by Deep Cogito from DeepSeek's DeepSeek-V3-Base, which Deep Cogito describes as the open-licensed DeepSeek base model from November 2024. The base model is not a catalog member and is not treated as U.S.-developed.147
- Derived from: DeepSeek-V3-Base (DeepSeek) (external site: huggingface.co) — Third-party base model; not a catalog member.
Other releases in the Cogito family
- Cogito v2 Preview Llama 405BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
Post-trained and published by Deep Cogito Inc., headquartered in San Francisco (see the deep-cogito record). Eligibility covers Deep Cogito's post-training only; the DeepSeek-V3 base model it starts from was developed by DeepSeek and is not treated as a U.S.-developed model.4897
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.