Independent project. Not a U.S. government website.

USASI
Model release

EmbeddingGemma 300M

Release in the Gemma family · version embeddinggemma-300m

Maintained by Google DeepMind1

EmbeddingGemma is a multilingual text embedding model in the Gemma family, built from Gemma 3 with T5Gemma initialization. Google's documentation gives its size as 308M parameters; the Hugging Face card calls it a 300M-parameter model. It accepts up to 2,048 input tokens and outputs 768-dimensional vectors, which can be truncated to 512, 256, or 128 dimensions through Matryoshka Representation Learning. Google's release page dates its release to September 4, 2025.134Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Sep 4, 2025

Availability and license

Overall availability

Public

Downloadable from Hugging Face after logging in and acknowledging Google's usage license; Google's overview also points to Kaggle and Vertex AI. The Hugging Face signals conflict: the gate text says requests are processed immediately, while the repository metadata sets the gate to manual mode, which Hugging Face's documentation describes as the author choosing which users to approve. This catalog has not confirmed which applies in practice.21310

Availability fact review: Oct 2, 2026

Availability is separate from permission: read the license before using or redistributing.

EmbeddingGemma is listed in the appendix of the Gemma Terms of Use (last modified April 1, 2026), so the separate Apache 2.0 license for Gemma 4 does not apply to it. The terms allow use, modification, and distribution subject to the Gemma Prohibited Use Policy, which they incorporate. Redistributors must pass on the use restrictions and a copy of the terms, mark modified files, and include a specified notice file. Model Derivatives include models trained by distillation or on synthetic outputs from Gemma. Google may restrict use it believes violates the terms and may terminate them on breach. California law governs.5Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Unknown — no complete fact-level rights review
code
Unknown — no complete fact-level rights review
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.2. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.2): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for EmbeddingGemma 300M
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicSafetensors weights on Hugging Face behind a license acknowledgment. The gate text says requests are processed immediately, but the repository metadata sets the gate to manual approval; the two have not been reconciled.2110
Inference codeIs code for running the model published?PublicThe model card documents inference with Sentence Transformers, using the Gemma 3 implementation in Hugging Face Transformers as the backbone.1
Training codeIs the code used to train the model published?UnknownThe card says training used JAX and ML Pathways; no published training code was found. Google's documentation links a fine-tuning guide.13
Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness.PartialThe card describes about 320 billion training tokens drawn from web documents in over 100 languages, code and technical documents, and synthetic and task-specific data, with CSAM, sensitive-data, and quality filtering. The data itself is not released, and this review does not judge the description complete.1
Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable.UnknownNot assessed.
Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient.UnknownNot assessed.
Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above.UnknownNot assessed.
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe paper describes encoder-decoder initialization, geometric embedding distillation, a spread-out regularizer, and merging of checkpoints trained on varied data mixtures. The card names TPUv5e hardware. This is not a complete recipe.61
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe model card reports benchmark results; evaluation code was not reviewed.1

What it is useful for

The model card lists retrieval, semantic similarity, classification, clustering, code retrieval, question answering, and fact verification. Google's documentation presents it for on-device use on phones, laptops, and tablets, including offline use and mobile-first retrieval-augmented generation together with Gemma 3n.13Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The model card says EmbeddingGemma activations do not support float16 and recommends float32 or bfloat16. It shows separate query and document encoding methods in Sentence Transformers, and task-specific prompts such as a code-retrieval query prefix.1

Organization context

Provenance and derivatives

Built by Google DeepMind from Gemma 3, with T5Gemma initialization, according to the model card.1

  • Derived from: Gemma 3 — Base model generation; see the Gemma family record.
  • Derived from: T5Gemma — Used for initialization; no separate catalog record.

Other releases in the Gemma family

  • Gemma 4 12B UnifiedModel-disclosure tier (USASI rubric v0.2): Open-weight
  • Gemma 4 26B A4BModel-disclosure tier (USASI rubric v0.2): Open-weight
  • Gemma 4 31BModel-disclosure tier (USASI rubric v0.2): Open-weight

Gemma family overview

What this catalog does not know

Unknown means the sources reviewed for this record do not document it. It is not evidence that something does not exist.
  • Training code: unknown.
  • Training-data access: unknown.
  • Complete training pipeline: unknown.
  • Legacy data assessment (v0.1): unknown.

Have a primary source? How to report a correction.

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: Documented U.S. control

The model card names Google DeepMind as the author, and the Gemma Terms of Use that govern it are issued by Google LLC. Google's CEO announced Google DeepMind in 2023 as a group combining DeepMind and Google Research's Brain team. Alphabet Inc.'s fiscal 2025 Form 10-K gives its principal executive offices in Mountain View, California, and Exhibit 21.01 lists Google LLC as a Delaware subsidiary.15789

Assessed Oct 1, 2026

Sources

  1. 1.
    google/embeddinggemma-300m model card (external site: huggingface.co)

    Google DeepMind (via Hugging Face) · Model card · accessed Oct 2, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    Hugging Face model metadata for google/embeddinggemma-300m (external site: huggingface.co)

    Hugging Face · Repository · accessed Oct 2, 2026 · evidence reviewed Oct 2, 2026

  3. 3.
    EmbeddingGemma model overview (external site: ai.google.dev)

    Google AI for Developers · Documentation · accessed Oct 2, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    Gemma releases (external site: ai.google.dev)

    Google AI for Developers · Release notes · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    Gemma Terms of Use (external site: ai.google.dev)

    Google AI for Developers · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    EmbeddingGemma: Powerful and Lightweight Text Representations (external site: arxiv.org)

    EmbeddingGemma Team, Google (via arXiv) · Paper · published Sep 24, 2025 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    Google DeepMind: Bringing together two world-class AI teams (external site: blog.google)

    Google · Announcement · published Apr 20, 2023 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  8. 8.
    Alphabet Inc. Form 10-K for fiscal 2025, cover page (XBRL viewer) (external site: sec.gov)

    Alphabet Inc. (via U.S. Securities and Exchange Commission EDGAR) · Filing · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  9. 9.
    Alphabet Inc. Form 10-K fiscal 2025, Exhibit 21.01 (subsidiaries) (external site: sec.gov)

    Alphabet Inc. (via U.S. Securities and Exchange Commission EDGAR) · Filing · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  10. 10.
    Gated models (Hugging Face Hub documentation) (external site: huggingface.co)

    Hugging Face · Documentation · accessed Oct 2, 2026 · evidence reviewed Oct 2, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project