Profluent-E1 300M (E1-300m)
Release in the Profluent-E1 family · version Profluent-Bio/E1-300m
The 300M-parameter member of the Profluent-E1 family of retrieval-augmented protein encoder models, released in November 2025 alongside E1-600m and E1-150m. The Hugging Face safetensors metadata counts 274,317,346 parameters stored in BF16. It can encode a single protein sequence or a query preceded by homologous sequences.1237
- Model hub: E1-300m weights (Hugging Face) (external site: huggingface.co)
- Repository: GitHub repository (Profluent-AI/E1) (external site: github.com)
- License: Profluent-E1 Clickthrough License Agreement (external site: github.com)
- Paper: Profluent-E1 preprint (bioRxiv 2025.11.12.688125) (external site: biorxiv.org)
Availability and license
Overall availability
Downloadable from Hugging Face without a gate, but the license states that downloading or using Profluent-E1 means accepting the Profluent-E1 Clickthrough License Agreement and its Attribution Guidelines. No biosecurity-specific access conditions were found in the README, license, or announcement.2437
Availability is separate from permission: read the license before using or redistributing.
Profluent-E1 Clickthrough License Agreement (external site: raw.githubusercontent.com)4612
Apache License 2.0 (model code used separately from the weights, per LICENSE section 4) (external site: raw.githubusercontent.com)43
The clickthrough agreement, offered by Profluent Bio Inc., grants a perpetual, worldwide, royalty-free copyright and patent license to use, modify, and redistribute Profluent-E1, with Apache-style redistribution conditions. It defines Profluent-E1 to include both the weights and the code, and says model code used separately from the weights is under Apache 2.0. The Attribution Guidelines require commercial entities to display "Profluent-E1" attribution and require any drug, target, hit, or lead created or identified using it to be identified as "Built with Profluent-E1", including in regulatory disclosures. Profluent may amend the agreement and may terminate it on breach; it is governed by California law with arbitration in Alameda County. The Hugging Face card lists the license as other (profluent-e1-clickthrough-license). One source file is adapted from flash-attention under BSD-3-Clause.46531
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Safetensors weights and config on Hugging Face without gating; use is subject to the clickthrough license.21 |
| Inference codeIs code for running the model published? | Public | The E1 GitHub repository provides the model code, batch preparation, and example notebooks.3 |
| Training codeIs the code used to train the model published? | Unknown | The README describes installation and inference only; no training code was found in the sources reviewed. |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The announcement says the models were trained on Profluent's Protein Atlas, described as billions of proteins, with homologous sequences supplied during training. This confirms partial disclosure, not completeness.7 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | No download of the Profluent Protein Atlas was found in the sources reviewed. |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | Not assessed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Unknown | The announcement describes the architecture at a high level; the preprint was not reviewed for this record. |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The README and announcement report ProteinGym and CAMEO results as charts, and the repository includes a zero-shot fitness prediction notebook; full evaluation scripts were not identified.37 |
What it is useful for
Run and use notes
- The README installs the code from the GitHub repository with pip. The models run on CPU or GPU; because they were trained in BF16 precision, the README recommends a GPU with BF16 support (CUDA capability 8.0 or higher) and optionally the flash-attn package.3
Organization context
Provenance and derivatives
Trained by Profluent on its Profluent Protein Atlas, which Profluent describes as a curated protein dataset; the dataset itself is not published in the sources reviewed.7
Other releases in the Profluent-E1 family
- Profluent-E1 150M (E1-150m)Model-disclosure tier (USASI rubric v0.2): Open-weight
- Profluent-E1 600M (E1-600m)Model-disclosure tier (USASI rubric v0.2): Open-weight
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
- Training recipe: unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.