USASI
Model release

V-JEPA 2.1 ViT-G/16 (384 px)

Release in the V-JEPA 2 family · version 2.1 ViT-G/16, 384 resolution (vjepa2_1_vit_gigantic_384)

Maintained by Meta (FAIR)16

The largest V-JEPA 2.1 model, a 2-billion-parameter ViT-G/16 video and image encoder released by Meta in March 2026. V-JEPA 2.1 changes the V-JEPA 2 recipe to learn dense, temporally consistent features, using a dense predictive loss over all tokens, self-supervision at several intermediate layers, and separate tokenizers for images and videos. Meta distilled this model into smaller ViT-B and ViT-L variants.16

Last reviewedEntry updated Documented release Mar 16, 2026

Availability and license

Overall availability

Public

Direct checkpoint download linked from the repository README, or loaded through PyTorch Hub as vjepa2_1_vit_gigantic_384. This review found no Hugging Face repository for V-JEPA 2.1.1

Availability is separate from permission: read the license before using or redistributing.

The repository README says most of the project is licensed under MIT, with three data-augmentation and worker-initialization files under Apache 2.0. It does not name a separate license for the V-JEPA 2.1 checkpoints, so no weights license is recorded here.1

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

No license for the weights is recorded in this catalog.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for V-JEPA 2.1 ViT-G/16 (384 px)
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicDirect download link in the README; no gating.1
Inference codeIs code for running the model published?PublicThe repository provides PyTorch Hub loaders for the V-JEPA 2.1 models.1
Training codeIs the code used to train the model published?PublicThe repository includes a V-JEPA 2.1 pretraining loop (app/vjepa_2_1) and pretraining and cooldown configurations for ViT-G/16.13
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe paper describes VisionMix-163M, which replaces V-JEPA 2's ImageNet images with LVD-142M and rebalances the video sources (SSv2, Kinetics, HowTo100M, YT-Temporal-1B). LVD-142M is a curated image set from Meta's earlier work that this review did not find released.6
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe paper describes the training phases (135,000 pretraining iterations, then a 12,000- iteration cooldown at higher resolution and more frames), and configs are published. The published ViT-G/16 cooldown config is named for 256-pixel input, while the checkpoint is a 384-pixel model; this review did not confirm that the configs reproduce the released checkpoint.631
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PublicThe repository publishes frozen-evaluation code and a V-JEPA 2.1 evaluation config set for this model (configs/eval_2_1/vitG-384), covering SSv2, Diving48, EPIC-KITCHENS-100, Kinetics-400, ImageNet-1k, COIN, and Jester. This review did not confirm configs for every task reported in the paper, such as depth estimation or robot grasping.14

What it is useful for

Dense and global visual features for tasks the paper evaluates, including depth estimation, action recognition and anticipation, robot grasping, and navigation.56

Organization context

Other releases in the V-JEPA 2 family

V-JEPA 2 family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

The paper lists all authors as affiliated with FAIR at Meta (the first author also lists Universidad de Zaragoza, with the work done at Meta), and the checkpoint and code are published by Meta in the facebookresearch vjepa2 repository under a Meta Platforms copyright. Meta Platforms, Inc. has its principal executive offices in Menlo Park, California, per its Form 10-K.6127

Assessed Sep 29, 2026

Sources

  1. 1.
    facebookresearch/vjepa2 README (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  2. 2.
    facebookresearch/vjepa2 LICENSE (MIT) (external site: github.com)

    Meta (GitHub) · License · accessed Sep 29, 2026

  3. 3.
    vjepa2 configs/train_2_1/vitG16 (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  4. 4.
    vjepa2 configs/eval_2_1/vitG-384 (external site: github.com)

    Meta (GitHub) · Repository · accessed Sep 29, 2026

  5. 5.
    V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (arXiv 2603.14482) (external site: arxiv.org)

    arXiv (Meta FAIR authors) · Paper · published Mar 15, 2026 · accessed Sep 29, 2026

  6. 6.
    V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (arXiv 2603.14482v3, HTML) (external site: arxiv.org)

    arXiv (Meta FAIR authors) · Paper · published Jun 11, 2026 · accessed Sep 29, 2026

  7. 7.
    Meta Platforms, Inc. Form 10-K for the fiscal year ended December 31, 2025 (external site: sec.gov)

    Meta Platforms, Inc. (U.S. SEC filing) · Filing · published Jan 29, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project