Zamba2-7B
Release in the Zamba family · version Zamba2-7B
The 7B base model in Zyphra's Zamba2 series, announced in October 2024. It interleaves Mamba2 blocks with two shared attention blocks in an alternating pattern and adds LoRA projectors to the shared MLP blocks so that each shared block can specialize by depth.21
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Release notes: Zamba2-7B announcement (external site: zyphra.com)
- Repository: Zamba2 standalone PyTorch implementation (external site: github.com)
- Documentation: Zamba2 in Hugging Face Transformers (external site: huggingface.co)
Availability and license
Overall availability
The Hugging Face repository is gated: users must agree to share their contact information before they can access the files. The Hugging Face metadata lists the gating mode as automatic approval, under which access is granted as soon as the request is sent.189
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Downloadable from Hugging Face after agreeing to share contact information; access is approved automatically.189 |
| Inference codeIs code for running the model published? | Public | Supported in Hugging Face Transformers from version 4.48.0; Zyphra also publishes a standalone PyTorch implementation of the Zamba2 models on GitHub.342 |
| Training codeIs the code used to train the model published? | Unknown | Zyphra's announcement says training used an internal framework built on Megatron-LM; this catalog found no public release of that framework.2 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card says pretraining used text and code from open web datasets, including Zyphra's Zyda (which is downloadable), followed by an annealing phase on a smaller high-quality mixture. No complete list of the mixture was found.16 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | Two-phase pretraining (main pretraining, then annealing) and the architecture are described at a high level in the model card and announcement.12 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Unknown | Not assessed. |
What it is useful for
Run and use notes
- The model card's quick start installs mamba-ssm (v2.1.0) and causal-conv1d for the optimized Mamba2 kernels and loads the model with Transformers on a CUDA device in bfloat16. It says running without those kernels is not recommended because latency and memory use are significantly higher.1
Organization context
Provenance and derivatives
A Zyphra base model pretrained on open web text and code, including Zyphra's Zyda dataset, using the Mistral v0.1 tokenizer. It serves as the language backbone of Zamba2-VL-7B.1367
- Derived from: Zyda (external site: huggingface.co) — Zyphra pretraining dataset named in the model card.
Catalog records that name this entry in their provenance:
Other releases in the Zamba family
- Zamba2-VL-7BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.