Phi-4-reasoning-vision-15B
Release in the Phi family · version 4-reasoning-vision-15B
Phi-4-reasoning-vision-15B is a 15-billion-parameter multimodal reasoning model from Microsoft. It combines the Phi-4-Reasoning language model with a SigLIP-2 vision encoder, takes text and images as input, produces text, and has a 16,384-token context length. It can either reason step by step or answer directly, depending on the task.1
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: GitHub repository (external site: github.com)
- Release notes: Microsoft Research blog post (external site: microsoft.com)
- Documentation: Data card (external site: huggingface.co)
Availability and license
Overall availability
Weights can be downloaded from Hugging Face without a gated access request. Microsoft also offers the model as a hosted deployment on Microsoft Foundry.21
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published on Hugging Face without gating.21 |
| Inference codeIs code for running the model published? | Public | The GitHub repository has Transformers and vLLM directories, and the model card documents the required torch, transformers, and vllm versions.41 |
| Training codeIs the code used to train the model published? | Unknown | The Microsoft Research blog says fine-tuning code was released with the model, but the GitHub repository contains only Transformers and vLLM inference code, and this review did not locate the fine-tuning code elsewhere.54 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card and blog describe about 200 billion tokens of multimodal data. Most of it comes from filtered open-source vision-language datasets, with added internal Microsoft domain data and targeted acquisitions. The blog says some responses were regenerated with GPT-4o and o4-mini. An EU-format data card is published; the training data itself is not.153 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The model card describes supervised fine-tuning on a mix of reasoning and non-reasoning data, the mid-fusion architecture, and compute (240 NVIDIA B200 GPUs for 4 days). A technical report is linked but was not reviewed here.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The model card reports results and names the open-source evaluation frameworks used (Eureka ML Insights and VLMEvalKit). Microsoft says benchmark logs were released; this review did not locate them.15 |
What it is useful for
The model card lists scientific and mathematical reasoning over visual inputs such as diagrams, charts, and documents, and computer-use agent tasks such as locating interface elements on screens. It also covers captioning, visual question answering, and OCR, and is trained primarily on English.1
Run and use notes
- The model card lists torch 2.7.1 or later and transformers 4.57.1 or later (vllm 0.15.2 or later for vLLM). It says the model was tested on NVIDIA A6000, A100, H100, and B200 GPUs under Ubuntu 22.04.5 and recommends serving it with vLLM in bf16 precision.1
Organization context
Provenance and derivatives
Built on Microsoft's Phi-4-Reasoning language model and a SigLIP-2 vision encoder, according to the model card. The Microsoft Research blog says some training responses were regenerated with GPT-4o and o4-mini.15
- Derived from: Phi-4-Reasoning (external site: huggingface.co) — Language model backbone listed as a model dependency.
- Derived from: SigLIP-2 — Vision encoder named in the model card.
Other releases in the Phi family
- Phi-4-mini-flash-reasoningModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
The model card names Microsoft Corporation as developer, with Microsoft Ireland Operations Limited as authorized representative for the EU. The EU-format data card lists the Irish entity in its developer field, but the release is a Microsoft product published on Microsoft's Hugging Face and GitHub accounts. Microsoft Corporation lists Redmond, Washington as its address on its Form 10-K cover page.136
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.