Hermes 4.3 36B
Release in the Hermes family · version Hermes-4.3-36B
Maintained by Nous Research13
A 36B-parameter hybrid reasoning model that Nous Research post-trained from ByteDance Seed's Seed-OSS-36B-Base. It was Nous's first Hermes model trained in a decentralized way over the internet on the Psyche network, with context extended up to 512K tokens. Nous also released a centrally trained version of the same model as a research comparison.134
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Release notes: Hermes 4.3 announcement (external site: nousresearch.com)
- Paper: Hermes 4 technical report (arXiv 2508.18255) (external site: arxiv.org)
- Dataset hub: Evaluation responses and scores (Hugging Face dataset) (external site: huggingface.co)
- Repository: Nous Research TorchTitan fork (external site: github.com)
Availability and license
Overall availability
Weights download from Hugging Face without an access gate; GGUF quantizations are in a separate repository. The model card also lists Nous Portal and Chutes as inference providers.1
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | BF16 safetensors in an ungated Hugging Face repository; the centrally trained comparison version is published separately.13 |
| Inference codeIs code for running the model published? | Public | The model card gives a Transformers example and says vLLM and SGLang include tool-call parsers for the Hermes format.1 |
| Training codeIs the code used to train the model published? | Partial | Nous says it trained Hermes 4.3 with its custom TorchTitan fork (BSD-3-Clause) and links the open-source Psyche code (Apache 2.0). This covers Nous's post-training only; the code used to pretrain the Seed-OSS base model is not part of this release.378 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card describes a post-training corpus of about 5 million samples mixing reasoning and non-reasoning data, and the Hermes 4 technical report describes how the Hermes 4 data was synthesized (DataForge, Atropos rejection sampling). The Hermes 4.3 post-training dataset itself was not found as a public release, and the base model's pretraining data is described only in ByteDance's materials.153 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The announcement describes decentralized training on Psyche using the DisTrO optimizer, spread across 24 Psyche nodes, with a larger training set and a run twice as large as Hermes 4; the Hermes 4 report documents the general post-training method. Full hyperparameters for this release were not found.35 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | Nous publishes the full evaluation responses and scores for this model (and for the centrally trained version) as Hugging Face datasets; the Hermes 4 report describes the evaluation harness, which uses lighteval, EQBench, and Atropos.365 |
What it is useful for
Nous presents it as a general assistant with optional reasoning, function calling and tool use, and schema-following JSON output.1
Run and use notes
- The model card recommends temperature 0.6, top_p 0.95, and top_k 20, uses the Llama 3 chat format with an optional thinking flag, and suggests tensor-parallel serving with SGLang or vLLM for multi-GPU nodes.1
Organization context
Provenance and derivatives
Fine-tuned by Nous Research from Seed-OSS-36B-Base, a base model developed and released by ByteDance's Seed Team under Apache 2.0. The base model is not a catalog member and is not treated as U.S.-developed; Nous Research's contribution is the post-training data and the decentralized post-training run on Psyche.139
- Derived from: Seed-OSS-36B-Base (ByteDance Seed Team) (external site: huggingface.co) — Third-party base model; not a catalog member.
Other releases in the Hermes family
- Hermes 4 70BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
This fine-tune was made and published by Nous Research, Inc., a Delaware corporation with U.S. addresses (see the nous-research record). Eligibility covers Nous Research's post-training only; the Seed-OSS-36B-Base model it starts from was developed by ByteDance's Seed Team and is not treated as a U.S.-developed model.1109
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.