USASI
Model release

Zamba2-7B

Release in the Zamba family · version Zamba2-7B

Maintained by Zyphra12

The 7B base model in Zyphra's Zamba2 series, announced in October 2024. It interleaves Mamba2 blocks with two shared attention blocks in an alternating pattern and adds LoRA projectors to the shared MLP blocks so that each shared block can specialize by depth.21

Last reviewedEntry updated Documented release Oct 14, 2024

Availability and license

Overall availability

Public

The Hugging Face repository is gated: users must agree to share their contact information before they can access the files. The Hugging Face metadata lists the gating mode as automatic approval, under which access is granted as soon as the request is sent.189

Availability is separate from permission: read the license before using or redistributing.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Zamba2-7B
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicDownloadable from Hugging Face after agreeing to share contact information; access is approved automatically.189
Inference codeIs code for running the model published?PublicSupported in Hugging Face Transformers from version 4.48.0; Zyphra also publishes a standalone PyTorch implementation of the Zamba2 models on GitHub.342
Training codeIs the code used to train the model published?UnknownZyphra's announcement says training used an internal framework built on Megatron-LM; this catalog found no public release of that framework.2
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe model card says pretraining used text and code from open web datasets, including Zyphra's Zyda (which is downloadable), followed by an annealing phase on a smaller high-quality mixture. No complete list of the mixture was found.16
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialTwo-phase pretraining (main pretraining, then annealing) and the architecture are described at a high level in the model card and announcement.12
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.UnknownNot assessed.

What it is useful for

The model card presents it as a general-purpose base model for efficient inference with low memory use on consumer hardware. It is not tuned for instruction following or chat and has no moderation; Zyphra released a separate instruction-tuned Zamba2-7B-Instruct.12

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The model card's quick start installs mamba-ssm (v2.1.0) and causal-conv1d for the optimized Mamba2 kernels and loads the model with Transformers on a CUDA device in bfloat16. It says running without those kernels is not recommended because latency and memory use are significantly higher.1

Organization context

Provenance and derivatives

A Zyphra base model pretrained on open web text and code, including Zyphra's Zyda dataset, using the Mistral v0.1 tokenizer. It serves as the language backbone of Zamba2-VL-7B.1367

Catalog records that name this entry in their provenance:

Other releases in the Zamba family

  • Zamba2-VL-7BModel-disclosure tier (USASI rubric v0.1): Open-weight

Zamba family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Developed and published by Zyphra, whose terms of use name Zyphra Technologies, Inc. at 415 Mission Street, San Francisco, and whose May 2026 announcement describes it as headquartered in San Francisco.11011

Assessed Sep 29, 2026

Sources

  1. 1.
    Zyphra/Zamba2-7B model card (external site: huggingface.co)

    Zyphra · Model card · accessed Sep 29, 2026

  2. 2.
    Zamba2-7B (external site: zyphra.com)

    Zyphra · Announcement · published Oct 14, 2024 · accessed Sep 29, 2026

  3. 3.
    Zamba2 | Hugging Face Transformers documentation (external site: huggingface.co)

    Hugging Face · Documentation · accessed Sep 29, 2026

  4. 4.
    Zyphra/Zamba2 README (external site: github.com)

    Zyphra · Repository · accessed Sep 29, 2026

  5. 5.
  6. 6.
    Zyphra/Zyda dataset card (external site: huggingface.co)

    Zyphra · Dataset card · accessed Sep 29, 2026

  7. 7.
    Zyphra/Zamba2-VL-7B model card (external site: huggingface.co)

    Zyphra · Model card · accessed Sep 29, 2026

  8. 8.
  9. 9.
    Gated models | Hugging Face Hub documentation (external site: huggingface.co)

    Hugging Face · Documentation · accessed Sep 29, 2026

  10. 10.
    Terms of Use | Zyphra (external site: zyphra.com)

    Zyphra Technologies, Inc. · Official page · published May 4, 2026 · accessed Sep 29, 2026

  11. 11.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project