Zamba2-VL-7B
Release in the Zamba family · version Zamba2-VL-7B
The largest of the three Zamba2-VL vision-language models Zyphra released in June 2026, with about 8B parameters. It pairs the Zamba2-7B hybrid state-space language model with the Qwen2.5-VL vision encoder and handles single- and multi-image understanding and grounding.21
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Model hub: Zamba2-VL collection (Hugging Face) (external site: huggingface.co)
- Release notes: Zamba2-VL announcement (external site: zyphra.com)
- Paper: Zamba2-VL technical report (arXiv 2606.00390) (external site: arxiv.org)
- Repository: Inference code (Zyphra Transformers fork, zamba2-vl branch) (external site: github.com)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.1
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | 1 |
| Inference codeIs code for running the model published? | Public | Inference uses the zamba2-vl branch of Zyphra's Transformers fork (based on Transformers 4.57.1), qwen-vl-utils, flash_attn, and Zyphra's forks of mamba-ssm and causal-conv1d, per the model card.12 |
| Training codeIs the code used to train the model published? | Unknown | The technical report points to the checkpoints and inference code; it does not mention a training code release.3 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The technical report names the public datasets used in its training stages, for example LLaVA-ReCap-558K, FineVision, PixMo, The Cauldron, and DocMatix, and says document and OCR data were upsampled.3 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The report describes three stages: alignment with only the MLP adapter trained (vision encoder and language backbone frozen), then joint pretraining and supervised fine-tuning of the full model.3 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The model card reports scores measured with Zyphra's own evaluation harness based on VLMEvalKit; the harness itself is not linked.1 |
What it is useful for
The model card describes it as a generalist model for on-device applications and reports results on visual question answering, OCR, chart and document understanding, and multimodal reasoning benchmarks. Zyphra's announcement says the hybrid backbone keeps a fixed-size recurrent state rather than a growing key-value cache, which lowers time to first token on long inputs.12
Run and use notes
- The model card says the model can run without the optimized Mamba2 kernels but that this is not recommended because latency and memory use are significantly higher. Its example loads the model in bfloat16 on a CUDA device with FlashAttention 2.1
Organization context
Provenance and derivatives
Built by Zyphra on its Zamba2-7B language model, with the vision transformer from Qwen2.5-VL as the vision backbone; Qwen is a model family built by Alibaba Cloud. The model uses the Mistral v0.1 tokenizer.126
- Derived from: Zamba2-7B — Language backbone.
- Derived from: Qwen2.5-VL vision encoder (external site: huggingface.co) — Vision backbone, developed by Alibaba Cloud's Qwen team. The model card and technical report do not say which Qwen2.5-VL checkpoint supplied the encoder.
Other releases in the Zamba family
- Zamba2-7BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
Developed and published by Zyphra. The technical report lists the authors' affiliation as Zyphra, San Francisco, CA, and Zyphra's terms of use name Zyphra Technologies, Inc. at a San Francisco address. The vision encoder comes from Alibaba Cloud's Qwen2.5-VL; that component is not thereby U.S.-developed, and eligibility here rests on Zyphra as the developer of this model.3516
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.