SynthID Text
Project record
Maintained by Google DeepMind13
SynthID Text is Google DeepMind's reference implementation of the text watermarking method described in its 2024 Nature paper. The method changes only how a language model samples tokens, using a configuration built around a sequence of integer keys, and detects the watermark from the text without running the model. The repository provides mean, weighted mean, and Bayesian scoring functions for detection, a mix-in that adds watermarking to Gemma and GPT-2 classes from Hugging Face Transformers, a Colab notebook, and human evaluation data.16
- Repository: Repository (external site: github.com)
- License: LICENSE (Apache-2.0) (external site: github.com)
- Website: Package on PyPI (external site: pypi.org)
- Paper: Paper (Nature) (external site: nature.com)
- Release notes: SynthID Text in Hugging Face Transformers (announcement) (external site: huggingface.co)
Availability and license
Overall availability
Source code and human evaluation data are public on GitHub, and the core library is published on PyPI as synthid-text.14
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (external site: github.com)21
Plain-language guide to the Apache License 2.0Creative Commons Attribution 4.0 International (materials other than software) (external site: github.com)1
Plain-language guide to the Creative Commons Attribution 4.0 International
The README states that all software is under Apache 2.0 and all other materials under CC-BY 4.0, with copyright held by DeepMind Technologies Limited, and that this is not an official Google product. The human evaluation data, which compares watermarked and unwatermarked text from Gemma 7B, falls under the "other materials" statement; the README gives code for loading the matching prompts separately from the ELI5 dataset.1
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Reviewed qualifying license recorded — check scope and conditions
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Python source for watermarking, the three detectors, and Bayesian detector training, with tests.1 |
| DocumentationIs user documentation published? | Public | The README explains the watermarking configuration, applying a watermark, detection, and the human data, and links a Colab notebook that runs end to end with Gemma or GPT-2.1 |
| InstallationAre installation instructions or packages publicly available? | Public | Installable from PyPI or from source with pip; the package requires Python 3.9 or later and pins specific versions of PyTorch (2.4.0), Transformers (4.43.3), and NumPy.13 |
| Supported platformsAre supported operating systems or hardware documented? | Partial | The package metadata lists Python 3.10 and 3.11 and "OS Independent", but it depends on jax[cuda]; the README describes running the notebook on Colab runtimes or locally from source. No tested platform list is given.31 |
| Release statusAre versioned releases published? | Public | Tagged releases 0.1 through 0.2.1 were published in November 2024, all marked as pre-releases on GitHub, and PyPI shows 0.2.1 (uploaded 2024-11-14) as the latest version, marked "Development Status :: 4 - Beta".54 |
What it is useful for
Reproducing and studying the paper's watermarking and detection methods. The README says the code is for reference and research reproducibility, is not intended for production use, and points to the SynthID Text implementation in Hugging Face Transformers for production.1
Run and use notes
- The README says the Bayesian detector must be trained on watermarked and unwatermarked text for each watermarking key, with training data representative of the text the system will generate. It also notes that the mix-in uses a static watermarking configuration and that the hashing function gives no guarantee of cryptographic security.1
- For the notebook, the README suggests a GPU with 16GB of memory (such as a T4) for Gemma v1.0 2B IT, 32GB (such as an A100) for Gemma v1.0 7B IT, and any runtime for GPT-2. Its watermarking example loads Gemma in bfloat16; the README does not say what precision the memory suggestions assume.1
- Google DeepMind and Hugging Face's announcement says SynthID Text was added to Transformers in version 4.46.0, that watermarking is less effective on factual responses, and that detection confidence falls when text is thoroughly rewritten or translated.7
Organization context
U.S. eligibility
Eligible · basis: Documented U.S. control
The repository is in Google DeepMind's GitHub organization, and its license section and package metadata name DeepMind Technologies Limited, a UK company, as copyright holder and author. UK Companies House lists DeepMind Holdings Limited as holding 75% or more of DeepMind Technologies Limited's shares and voting rights, and lists Alphabet Inc., a Delaware corporation, as holding 75% or more of DeepMind Holdings Limited's. Alphabet's fiscal 2025 Form 10-K gives its principal executive offices in Mountain View, California. This is the same us-control basis as the catalog's google-deepmind record, which treats Google DeepMind as part of Google: Google's CEO announced it in April 2023 as a unit combining DeepMind and Google Research's Brain team. This review did not examine which entities employ the repository's maintainers.13891011
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.