FastVLM 7B
Release in the FastVLM family · version FastVLM-7B
FastVLM 7B is the 7B variant of Apple's FastVLM release. It combines Apple's FastViTHD vision encoder with the Qwen2-7B language model in a LLaVA-style architecture and accepts an image plus a text prompt to generate text.173
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: ml-fastvlm repository (external site: github.com)
- Paper: FastVLM: Efficient Vision Encoding for Vision Language Models (arXiv 2412.13303) (external site: arxiv.org)
- License: License (Apple Machine Learning Research Model License) (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without an access gate and from Apple's CDN via the repository's download script, under a license that permits only non-commercial research use.274
Availability is separate from permission: read the license before using or redistributing.
Apple Machine Learning Research Model License Agreement (external site: huggingface.co)426
Apple software license (ml-fastvlm code) (external site: raw.githubusercontent.com)57
The weights license grants a personal, non-exclusive, revocable license to use, copy, modify, distribute, and create derivatives only for non-commercial scientific research, excluding product development and commercial use; derivatives are held to the same limit. Redistribution requires passing on the agreement with a specified attribution notice. No patent rights are granted, and the agreement is governed by California law. The Qwen2-7B language model it builds on is separately published under Apache 2.0 on Hugging Face.410
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Weights are published on Hugging Face without an access gate; the repository also links stage 2 and stage 3 PyTorch checkpoints and an int4 Apple silicon export.27 |
| Inference codeIs code for running the model published? | Public | The ml-fastvlm repository provides a predict.py script, export code for Apple silicon, and an iOS/macOS demo app; the card also shows Transformers inference with remote code.71 |
| Training codeIs the code used to train the model published? | Unknown | The README says FastVLM variants were trained with the LLaVA codebase and directs users to LLaVA's instructions for training; FastVLM-specific training scripts were not found.7 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The paper describes the datasets used in each training stage (for example the LLaVA-558K alignment and LLaVA-665K instruction data in its base setup) and the DataCompDR-1B data used to pretrain FastViTHD; availability of every dataset for the released checkpoint was not verified.8 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The paper describes the multi-stage training setup (connector training, resolution scaling, visual instruction tuning, and a further stage 3), with some hyperparameters and compute details.8 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card and paper report benchmark results; evaluation code for re-running them was not identified.18 |
What it is useful for
Run and use notes
- The repository recommends exporting PyTorch checkpoints for Apple silicon with its model_export code and provides a pre-exported 7B stage 3 checkpoint with int4 quantization for the language model.7
Organization context
Provenance and derivatives
Built by Apple from its FastViTHD vision encoder and the Qwen2-7B language model, which was developed by the Qwen Team at Alibaba Group (per the Qwen2 Technical Report) and is published on Hugging Face. The configuration identifies the architecture as LlavaQwen2ForCausalLM.731011
- Derived from: Qwen2-7B (external site: huggingface.co) — Language model decoder.
Other releases in the FastVLM family
No other releases in this family have been assessed.
U.S. eligibility
Eligible · basis: U.S. headquarters
The model license states the model is developed and released by Apple Inc., and the paper lists Apple as the authors' affiliation. Apple Inc. has its principal executive offices in Cupertino, California, per its Form 10-K for fiscal year 2025. The language model is the third-party Qwen2-7B, recorded under provenance; that base is not U.S.-developed.48127
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.