Florence-2-large
Release in the Microsoft Florence-2 family · version Florence-2-large
Maintained by Microsoft (Azure AI)15
Florence-2-large is the larger pretrained Florence-2 vision model, listed at 0.77B parameters on its model card. It takes an image and a task prompt and returns text, boxes, or polygons for tasks such as captioning, detection, grounding, and OCR. Since a December 2024 update, the Hugging Face repository holds a version continued-pretrained for a 4k context length on 0.1B samples, which the card says might not be trained well; the same update changed OCR output to include line separators.14
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- License: LICENSE (MIT) (external site: huggingface.co)
- Paper: Florence-2 technical report (arXiv 2311.06242) (external site: arxiv.org)
- Documentation: Sample inference notebook (external site: huggingface.co)
Availability and license
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Safetensors and PyTorch weight files are in the Hugging Face repository, which is not gated.3 |
| Inference codeIs code for running the model published? | Public | The repository includes the modeling and processing code, and the card gives Transformers examples for each task plus a sample inference notebook.13 |
| Training codeIs the code used to train the model published? | Unknown | No official Florence-2 training code was linked from the model card, the report, or the Microsoft Research publication page.156 |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The report describes building FLD-5B from images drawn from ImageNet-22k, Object 365, Open Images, Conceptual Captions, and LAION, annotated by specialist models and refined through filtering and iterative model refinement. This review confirms disclosure, not completeness across every provenance, labeling, and processing requirement.5 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | No official download of the FLD-5B annotations was linked from the model card, the report, or the Microsoft Research publication page.156 |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The report gives the optimizer (AdamW with cosine decay), maximum learning rate, warmup, batch size, image sizes, sample counts, and initialization for the original large model; the card describes the later 4k-context continued pretraining only briefly.51 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card and report publish zero-shot and fine-tuned results; evaluation code was not found.15 |
What it is useful for
Run and use notes
- The model card documents loading the model with Hugging Face Transformers (AutoModelForCausalLM with trust_remote_code) and states that all models were trained with float16; its examples use float16 on CUDA and float32 on CPU.1
Organization context
Provenance and derivatives
Trained by Microsoft. The technical report says the image encoder and the multimodal encoder-decoder were initialized from UniCL and BART weights, respectively. The current repository weights are a continued-pretrained 4k-context version uploaded in December 2024.514
- Derived from: UniCL — Initialization for the DaViT image encoder, per the technical report.
- Derived from: BART — Initialization for the multimodal encoder-decoder, per the technical report.
Other releases in the Microsoft Florence-2 family
No other releases in this family have been assessed.
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Eligible · basis: U.S. headquarters
The technical report lists its authors under Azure AI, Microsoft, and the checkpoint is published in Microsoft's Hugging Face organization with an MIT license naming Microsoft Corporation. Microsoft Corporation lists One Microsoft Way, Redmond, Washington as its principal executive offices on the cover page of its Form 10-K for fiscal year 2026.5127
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.