EmbeddingGemma 2
Release in the Gemma family · version embeddinggemma-2
Maintained by Google DeepMind13
EmbeddingGemma 2 is a multimodal embedding model from Google DeepMind that maps text (including code), images, video, and audio, alone or combined in one input, into a single 768-dimensional vector space. The model card gives 740M parameters in total: a 270M-parameter text model plus a 170M vision encoder and a 300M audio encoder that can be loaded separately. All inputs share an 8,192-token context, and outputs can be shortened to 512, 256, or 128 dimensions. Google's release page dates it to October 6, 2026.135
- Model hub: Hugging Face model card (external site: huggingface.co)
- Documentation: EmbeddingGemma 2 model card (Google AI for Developers) (external site: ai.google.dev)
- Documentation: EmbeddingGemma overview (external site: ai.google.dev)
- License: Apache License 2.0 (Google's Gemma 4 license page) (external site: ai.google.dev)
- Release notes: Launch post (external site: blog.google)
Availability and license
Overall availability
Downloadable from Hugging Face without an access gate. Google's launch post says the weights are also published on Kaggle.26
Availability is separate from permission: read the license before using or redistributing.
The model card and the repository metadata give Apache 2.0, and the card links to Google's Gemma 4 license page, which reproduces the Apache License 2.0 text; the Hugging Face repository has no separate LICENSE file. The Gemma Terms of Use (last modified April 1, 2026) list the original EmbeddingGemma in their appendix but not EmbeddingGemma 2. The model card's ethics section separately says deployments must adhere to the Gemma Prohibited Use Policy; the Apache license text itself does not refer to that policy.1278
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Unknown — no complete fact-level rights review
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | A single BF16 safetensors file in the ungated Hugging Face repository; the repository metadata counts 744,371,512 parameters.2 |
| Inference codeIs code for running the model published? | Public | The model card documents inference with the Sentence Transformers library on top of Hugging Face Transformers. Google's launch post also names MLX, vLLM, llama.cpp, SGLang, Ollama, and LM Studio for serving the model.16 |
| Training codeIs the code used to train the model published? | Unknown | No training code was found. Google's EmbeddingGemma overview links a fine-tuning guide for EmbeddingGemma that uses Sentence Transformers.4 |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The card describes web documents in over 140 languages, code, images, video, audio, and paired cross-modal samples, with a January 2025 cutoff, and CSAM, sensitive-data, and quality and safety filtering. It gives no data sizes, the data is not released, and this review does not judge the description complete.1 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | Not assessed. |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Unknown | The card gives the architecture (24 layers, mean pooling, a 512-to-768 projection) but not the training procedure, and no technical report for this release was found.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card reports results on MTEB (multilingual and code), MIEB, MMEB v2, MSEB, and MAEB, including at truncated dimensions; evaluation code and settings were not reviewed.1 |
What it is useful for
The model card lists retrieval, classification, clustering, semantic similarity, and fact verification across text, code, images, video, and audio, for example search and retrieval-augmented generation over documents, code, or audio archives. It describes the model as designed to run on consumer hardware such as mobile devices and laptops.1
Run and use notes
- The model card says to run inference in bfloat16 or float32, not float16, because the model's activations exceed float16's range and can produce NaN or degraded embeddings without an error.1
- Text inputs take short task prefixes (for example "task: search result | query:" for queries and "title: ... | text:" for documents); images, video, and audio take none. The card's default budgets are 280 tokens per image, 140 per video frame (sampled at one frame per second), and 25 per second of 16 kHz mono audio, all within the shared 8,192-token context.1
- The vision and audio encoders can be left unloaded through Sentence Transformers' config_kwargs; the card lists an effective size of 270M parameters for text only and 440M for text and images. After truncating an embedding, the card says to L2-normalize it again, and queries and documents must use the same dimension.1
Organization context
Provenance and derivatives
Google describes EmbeddingGemma 2 as built on the Gemma 4 architecture. Its overview compares it with the original 308M-parameter text-only EmbeddingGemma, which it calls EmbeddingGemma 1. The card does not say whether EmbeddingGemma 2 was initialized from Gemma 4 weights.164
- Derived from: Gemma 4 — Architecture basis named by Google; see the Gemma family record.
- Derived from: EmbeddingGemma (300M) — Earlier text-only release that Google's overview calls EmbeddingGemma 1.
Other releases in the Gemma family
- EmbeddingGemma 300MModel-disclosure tier (USASI rubric v0.2): Open-weight
- Gemma 4 12B UnifiedModel-disclosure tier (USASI rubric v0.2): Open-weight
- Gemma 4 26B A4BModel-disclosure tier (USASI rubric v0.2): Open-weight
- Gemma 4 31BModel-disclosure tier (USASI rubric v0.2): Open-weight
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
- Training recipe: unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Eligible · basis: Documented U.S. control
The model card names Google DeepMind as the author. Google's CEO announced Google DeepMind in 2023 as a group within Google combining DeepMind and Google Research's Brain team. Alphabet Inc.'s fiscal 2025 Form 10-K gives its principal executive offices in Mountain View, California, and Exhibit 21.01 lists Google LLC as a Delaware subsidiary.191011
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.