EmbeddingGemma 300M
Release in the Gemma family · version embeddinggemma-300m
Maintained by Google DeepMind1
EmbeddingGemma is a multilingual text embedding model in the Gemma family, built from Gemma 3 with T5Gemma initialization. Google's documentation gives its size as 308M parameters; the Hugging Face card calls it a 300M-parameter model. It accepts up to 2,048 input tokens and outputs 768-dimensional vectors, which can be truncated to 512, 256, or 128 dimensions through Matryoshka Representation Learning. Google's release page dates its release to September 4, 2025.134
- Model hub: Hugging Face model card (external site: huggingface.co)
- Documentation: EmbeddingGemma overview (external site: ai.google.dev)
- License: Gemma Terms of Use (external site: ai.google.dev)
- Paper: EmbeddingGemma: Powerful and Lightweight Text Representations (arXiv) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable from Hugging Face after logging in and acknowledging Google's usage license; Google's overview also points to Kaggle and Vertex AI. The Hugging Face signals conflict: the gate text says requests are processed immediately, while the repository metadata sets the gate to manual mode, which Hugging Face's documentation describes as the author choosing which users to approve. This catalog has not confirmed which applies in practice.21310
Availability is separate from permission: read the license before using or redistributing.
EmbeddingGemma is listed in the appendix of the Gemma Terms of Use (last modified April 1, 2026), so the separate Apache 2.0 license for Gemma 4 does not apply to it. The terms allow use, modification, and distribution subject to the Gemma Prohibited Use Policy, which they incorporate. Redistributors must pass on the use restrictions and a copy of the terms, mark modified files, and include a specified notice file. Model Derivatives include models trained by distillation or on synthetic outputs from Gemma. Google may restrict use it believes violates the terms and may terminate them on breach. California law governs.5
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Unknown — no complete fact-level rights review
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Safetensors weights on Hugging Face behind a license acknowledgment. The gate text says requests are processed immediately, but the repository metadata sets the gate to manual approval; the two have not been reconciled.2110 |
| Inference codeIs code for running the model published? | Public | The model card documents inference with Sentence Transformers, using the Gemma 3 implementation in Hugging Face Transformers as the backbone.1 |
| Training codeIs the code used to train the model published? | Unknown | The card says training used JAX and ML Pathways; no published training code was found. Google's documentation links a fine-tuning guide.13 |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The card describes about 320 billion training tokens drawn from web documents in over 100 languages, code and technical documents, and synthetic and task-specific data, with CSAM, sensitive-data, and quality filtering. The data itself is not released, and this review does not judge the description complete.1 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | Not assessed. |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The paper describes encoder-decoder initialization, geometric embedding distillation, a spread-out regularizer, and merging of checkpoints trained on varied data mixtures. The card names TPUv5e hardware. This is not a complete recipe.61 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The model card reports benchmark results; evaluation code was not reviewed.1 |
What it is useful for
The model card lists retrieval, semantic similarity, classification, clustering, code retrieval, question answering, and fact verification. Google's documentation presents it for on-device use on phones, laptops, and tablets, including offline use and mobile-first retrieval-augmented generation together with Gemma 3n.13
Run and use notes
- The model card says EmbeddingGemma activations do not support float16 and recommends float32 or bfloat16. It shows separate query and document encoding methods in Sentence Transformers, and task-specific prompts such as a code-retrieval query prefix.1
Organization context
Provenance and derivatives
Other releases in the Gemma family
- Gemma 4 12B UnifiedModel-disclosure tier (USASI rubric v0.2): Open-weight
- Gemma 4 26B A4BModel-disclosure tier (USASI rubric v0.2): Open-weight
- Gemma 4 31BModel-disclosure tier (USASI rubric v0.2): Open-weight
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Eligible · basis: Documented U.S. control
The model card names Google DeepMind as the author, and the Gemma Terms of Use that govern it are issued by Google LLC. Google's CEO announced Google DeepMind in 2023 as a group combining DeepMind and Google Research's Brain team. Alphabet Inc.'s fiscal 2025 Form 10-K gives its principal executive offices in Mountain View, California, and Exhibit 21.01 lists Google LLC as a Delaware subsidiary.15789
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.