USASI
Model release

Zamba2-VL-7B

Release in the Zamba family · version Zamba2-VL-7B

Maintained by Zyphra12

The largest of the three Zamba2-VL vision-language models Zyphra released in June 2026, with about 8B parameters. It pairs the Zamba2-7B hybrid state-space language model with the Qwen2.5-VL vision encoder and handles single- and multi-image understanding and grounding.21

Last reviewedEntry updated Documented release Jun 2, 2026

Availability and license

Overall availability

Public

Downloadable from Hugging Face without gating.1

Availability is separate from permission: read the license before using or redistributing.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Zamba2-VL-7B
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?Public1
Inference codeIs code for running the model published?PublicInference uses the zamba2-vl branch of Zyphra's Transformers fork (based on Transformers 4.57.1), qwen-vl-utils, flash_attn, and Zyphra's forks of mamba-ssm and causal-conv1d, per the model card.12
Training codeIs the code used to train the model published?UnknownThe technical report points to the checkpoints and inference code; it does not mention a training code release.3
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe technical report names the public datasets used in its training stages, for example LLaVA-ReCap-558K, FineVision, PixMo, The Cauldron, and DocMatix, and says document and OCR data were upsampled.3
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe report describes three stages: alignment with only the MLP adapter trained (vision encoder and language backbone frozen), then joint pretraining and supervised fine-tuning of the full model.3
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe model card reports scores measured with Zyphra's own evaluation harness based on VLMEvalKit; the harness itself is not linked.1

What it is useful for

The model card describes it as a generalist model for on-device applications and reports results on visual question answering, OCR, chart and document understanding, and multimodal reasoning benchmarks. Zyphra's announcement says the hybrid backbone keeps a fixed-size recurrent state rather than a growing key-value cache, which lowers time to first token on long inputs.12

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The model card says the model can run without the optimized Mamba2 kernels but that this is not recommended because latency and memory use are significantly higher. Its example loads the model in bfloat16 on a CUDA device with FlashAttention 2.1

Organization context

Provenance and derivatives

Built by Zyphra on its Zamba2-7B language model, with the vision transformer from Qwen2.5-VL as the vision backbone; Qwen is a model family built by Alibaba Cloud. The model uses the Mistral v0.1 tokenizer.126

Other releases in the Zamba family

  • Zamba2-7BModel-disclosure tier (USASI rubric v0.1): Open-weight

Zamba family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

Developed and published by Zyphra. The technical report lists the authors' affiliation as Zyphra, San Francisco, CA, and Zyphra's terms of use name Zyphra Technologies, Inc. at a San Francisco address. The vision encoder comes from Alibaba Cloud's Qwen2.5-VL; that component is not thereby U.S.-developed, and eligibility here rests on Zyphra as the developer of this model.3516

Assessed Sep 29, 2026

Sources

  1. 1.
    Zyphra/Zamba2-VL-7B model card (external site: huggingface.co)

    Zyphra · Model card · accessed Sep 29, 2026

  2. 2.
    Zamba2-VL: Hybrid SSM Vision-Language Models (external site: zyphra.com)

    Zyphra · Announcement · published Jun 2, 2026 · accessed Sep 29, 2026

  3. 3.
    Zamba2-VL Technical Report (arXiv 2606.00390v1) (external site: arxiv.org)

    Zyphra · Paper · published May 29, 2026 · accessed Sep 29, 2026

  4. 4.
    Zyphra/transformers LICENSE (zamba2-vl branch) (external site: raw.githubusercontent.com)

    Hugging Face (fork maintained by Zyphra) · License · accessed Sep 29, 2026

  5. 5.
    Terms of Use | Zyphra (external site: zyphra.com)

    Zyphra Technologies, Inc. · Official page · published May 4, 2026 · accessed Sep 29, 2026

  6. 6.
    Qwen on Hugging Face (external site: huggingface.co)

    Qwen (Alibaba Cloud) · Repository · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project