SynthID Bio
Project record
Maintained by Google DeepMind12
SynthID Bio is a set of methods from Google DeepMind for watermarking AI-generated protein sequences and biomolecular structures so that their AI origin can later be detected. SynthID Bio-sequence adds watermarked sampling to the ProteinMPNN inverse-folding model, and SynthID Bio-structure is a fine-tuned AlphaFold 3 model that embeds a watermark in generated structures. The public repository contains the sequence watermarking code, a standalone detection script, and data supporting the paper.16
- Repository: Repository (external site: github.com)
- License: LICENSE (Apache-2.0) (external site: github.com)
- Paper: Paper (Nature) (external site: nature.com)
Availability and license
Overall availability
The sequence watermarking code, bundled ProteinMPNN parameters, and paper data are on GitHub. For the structure watermarking model, the README refers users to the AlphaFold 3 repository for downloading weights, which are governed by the AlphaFold 3 Model Parameters Terms of Use; this catalog did not confirm how the fine-tuned SynthID Bio-structure parameters are obtained.1
Availability is separate from permission: read the license before using or redistributing.
MIT License (bundled ProteinMPNN model parameters) (external site: github.com)13
Creative Commons Attribution 4.0 International (external site: github.com)14
The README states that all software is under Apache 2.0, that the ProteinMPNN model parameters are redistributed under their original MIT License, and that the data directory is under CC-BY 4.0. AlphaFold 3 model parameters are covered by the separate AlphaFold 3 Model Parameters Terms of Use, which the repository references rather than reproduces. The README also states that software, parameters, and outputs are for theoretical modeling only and not for clinical use.15
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Reviewed qualifying license recorded — check scope and conditions
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Public on GitHub, including a modified copy of ProteinMPNN with watermarking added (the README lists the changes and notes that ProteinMPNN's training code was removed) and a vendored copy of SynthID Text.1 |
| DocumentationIs user documentation published? | Public | The README documents setup, watermarked and unwatermarked generation examples, and detection.1 |
| InstallationAre installation instructions or packages publicly available? | Public | The README gives steps using uv or venv with Python 3.9, a requirements file, and an editable install of the bundled SynthID Text package.1 |
| Supported platformsAre supported operating systems or hardware documented? | Partial | The README's setup steps are written for Linux; other platforms are not documented.1 |
| Release statusAre versioned releases published? | Not public | The repository listed no releases or tags at review; the code is used from the main branch. It was published alongside the paper (Nature, online 2026-09-30).16 |
What it is useful for
Organization context
Provenance and derivatives
SynthID Bio-sequence modifies ProteinMPNN, an existing open-source inverse-folding model, and uses Google DeepMind's SynthID Text watermarking package. SynthID Bio-structure is a fine-tune of AlphaFold 3.16
- Derived from: ProteinMPNN — Vendored and modified in synthidbio_sequence/third_party/ProteinMPNN (MIT License).
- Derived from: SynthID Text — Vendored in synthidbio_sequence/third_party/synthid-text (Apache-2.0).
- Derived from: AlphaFold 3 — Base model for SynthID Bio-structure.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.