USASI
Model release

CLIP ViT-L/14

Release in the CLIP family · version ViT-L/14

Maintained by OpenAI31

CLIP ViT-L/14 pairs a ViT-L/14 Vision Transformer image encoder with a masked self-attention Transformer text encoder, trained contrastively to match images with their captions. OpenAI released it in January 2022; a variant further trained at 336-pixel resolution (ViT-L/14@336px) followed in April 2022.139

Last reviewedEntry updated Documented release Jan 2022

Availability and license

Overall availability

Public

Downloadable without gating from Hugging Face, and fetched by the openai/CLIP package's clip.load("ViT-L/14"). No license for the weights is stated (see license notes).26

Availability is separate from permission: read the license before using or redistributing.

The repository's MIT license covers "the Software" (the code). This review found no license statement for the model weights in the repository README, the model card, or the Hugging Face card, whose metadata carries no license field. Separately from licensing, the model card places all deployed uses out of scope.5432

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

No license for the weights is recorded in this catalog.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for CLIP ViT-L/14
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicPublished on Hugging Face without gating and downloadable through clip.load().26
Inference codeIs code for running the model published?PublicThe openai/CLIP package loads the model and encodes images and text; the Hugging Face card documents use with Transformers.41
Training codeIs the code used to train the model published?UnknownThe repository provides loading, inference, and evaluation examples; this review found no published training code.4
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe paper describes a dataset of 400 million image-text pairs gathered from public internet sources using 500,000 search queries. The model card says OpenAI will not release the dataset; only a list identifying a YFCC100M subset used in an ablation is published.938
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe paper describes training for 32 epochs with Adam, decoupled weight decay, a cosine schedule, and a 32,768 minibatch, and says the largest Vision Transformer took 12 days on 256 V100 GPUs. No training configuration files are published.9
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PublicThe repository publishes the class names and prompt templates used for the paper's zero-shot results, an ImageNet prompt-engineering notebook, and a linear-probe example.74

What it is useful for

Zero-shot image classification and image-text similarity for research. The model card says any deployed use case, commercial or not, is currently out of scope, surveillance and facial recognition are always out of scope, and use should be limited to English.13

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README documents installation with PyTorch 1.7.1 or later and torchvision, then loading a model by name with clip.load(); the Hugging Face card documents use through Transformers' CLIPModel and CLIPProcessor.41

Organization context

Other releases in the CLIP family

No other releases in this family have been assessed.

CLIP family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

The model card states that CLIP was developed by researchers at OpenAI, and the checkpoint is published by OpenAI in its openai/CLIP repository and Hugging Face account. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.3610

Assessed Sep 29, 2026

Sources

  1. 1.
    openai/clip-vit-large-patch14 model card (external site: huggingface.co)

    OpenAI (Hugging Face) · Model card · accessed Sep 29, 2026

  2. 2.
  3. 3.
    Model Card: CLIP (external site: github.com)

    OpenAI (GitHub) · Model card · accessed Sep 29, 2026

  4. 4.
    openai/CLIP README (external site: github.com)

    OpenAI (GitHub) · Repository · accessed Sep 29, 2026

  5. 5.
    openai/CLIP LICENSE (external site: github.com)

    OpenAI (GitHub) · License · accessed Sep 29, 2026

  6. 6.
    openai/CLIP clip/clip.py (checkpoint download URLs) (external site: github.com)

    OpenAI (GitHub) · Repository · accessed Sep 29, 2026

  7. 7.
  8. 8.
    openai/CLIP data/yfcc100m.md (YFCC100M subset) (external site: github.com)

    OpenAI (GitHub) · Documentation · accessed Sep 29, 2026

  9. 9.
    Learning Transferable Visual Models From Natural Language Supervision (arXiv 2103.00020, PDF) (external site: arxiv.org)

    arXiv (OpenAI authors) · Paper · published Feb 26, 2021 · accessed Sep 29, 2026

  10. 10.
    Exhibit 10.1: Equity commitment letter agreement between OpenAI Group PBC and Amazon (external site: sec.gov)

    U.S. Securities and Exchange Commission (Amazon.com, Inc. filing) · Filing · published Feb 27, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project