USASI
Model release

V-JEPA 2 ViT-g/16 (384 px)

Release in the V-JEPA 2 family · version 2 ViT-g/16, 384 resolution (vjepa2-vitg-fpc64-384)

Maintained by Meta (FAIR)31

The 1-billion-parameter ViT-g/16 encoder from Meta's V-JEPA 2, in the version that takes 64-frame clips at 384-pixel resolution. It was pretrained with self-supervision on VideoMix22M, a mix of about 22 million video and image samples. Meta's action-conditioned V-JEPA 2-AC world model was post-trained from a ViT-g/16 V-JEPA 2 encoder.137

Last reviewedEntry updated Documented release Jun 11, 2025

Availability and license

Overall availability

Public

Downloadable without gating from Hugging Face and as a direct checkpoint link in the repository README.23

Availability is separate from permission: read the license before using or redistributing.

The repository README says most of V-JEPA 2 is licensed under MIT, with three data-augmentation and worker-initialization files under Apache 2.0, but does not name a separate license for the checkpoints linked from the README. The Hugging Face card for this checkpoint declares Apache-2.0. Meta's announcement says the code and checkpoints are available for commercial and research applications.318

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-stack

Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for V-JEPA 2 ViT-g/16 (384 px)
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicUngated Hugging Face repository and direct download link.23
Inference codeIs code for running the model published?PublicThe repository provides PyTorch Hub loaders and a demo; the model card documents use with Hugging Face Transformers.31
Training codeIs the code used to train the model published?PublicThe repository publishes the V-JEPA 2 pretraining loop and launch commands, with pretraining and cooldown configurations for ViT-g/16 (including a 384-pixel cooldown).35
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe paper lists VideoMix22M's components (Something-Something v2, Kinetics, HowTo100M, YT-Temporal-1B, and ImageNet) and says all sources are publicly available. The YT-Temporal-1B portion was filtered by retrieval-based curation; this review did not find the curated subset list published.7
Training recipeAre the training configuration and procedure documented in enough detail to follow?PublicPretraining and cooldown configuration files are published, and the paper documents pretraining hyperparameters in its appendix.57
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PublicThe repository publishes attentive-probe evaluation code, evaluation configs for the ViT-g 384-pixel model (for example SSv2, Diving48, EPIC-KITCHENS-100, Kinetics-400, ImageNet-1k), and trained probe checkpoints.36

What it is useful for

The model card describes use as a video and image encoder for video classification, retrieval, and as the video encoder of vision-language models.1

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README notes that the code depends on decord, which does not support macOS, and documents Python 3.12 environment setup; the model card documents loading with Transformers' AutoModel and AutoVideoProcessor.31

Organization context

Other releases in the V-JEPA 2 family

V-JEPA 2 family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

The model card identifies FAIR at Meta as the developer, and the code is published by Meta in the facebookresearch GitHub organization. Meta Platforms, Inc. has its principal executive offices in Menlo Park, California, per its Form 10-K.139

Assessed Sep 29, 2026

Sources

  1. 1.
    facebook/vjepa2-vitg-fpc64-384 model card (external site: huggingface.co)

    Meta (Hugging Face) · Model card · accessed Sep 29, 2026

  2. 2.
  3. 3.
    facebookresearch/vjepa2 README (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  4. 4.
    facebookresearch/vjepa2 LICENSE (MIT) (external site: github.com)

    Meta (GitHub) · License · accessed Sep 29, 2026

  5. 5.
    vjepa2 configs/train/vitg16 (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  6. 6.
    vjepa2 configs/eval/vitg-384 (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  7. 7.
  8. 8.
    V-JEPA 2 announcement (Meta AI blog) (external site: ai.meta.com)

    Meta · Announcement · published Jun 11, 2025 · accessed Sep 29, 2026

  9. 9.
    Meta Platforms, Inc. Form 10-K for the fiscal year ended December 31, 2025 (external site: sec.gov)

    Meta Platforms, Inc. (U.S. SEC filing) · Filing · published Jan 29, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project