ether0 (24B)
Release in the ether0 family · version futurehouse/ether0
Maintained by FutureHouse123
The only published ether0 checkpoint: a 24B-parameter language model that reasons in English and outputs molecules as SMILES. It is fine-tuned and reinforcement-learning trained from Mistral-Small-24B-Instruct-2501, followed by safety post-training. It was posted to Hugging Face on June 4, 2025.12
- Model hub: Model card and weights (Hugging Face) (external site: huggingface.co)
- Repository: Reward functions and utilities (GitHub) (external site: github.com)
- Dataset hub: ether0 benchmark test set (Hugging Face) (external site: huggingface.co)
- Paper: Training a Scientific Reasoning Model for Chemistry (arXiv 2506.17238) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable from Hugging Face without gating. The card documents safety post-training (refusals for compounds on OPCW Schedules 1 and 2 and for requests such as making explosives or poisons) rather than any access restriction.12
Availability is separate from permission: read the license before using or redistributing.
The model card's license field is apache-2.0, and its licensing section calls the repository open weights under Apache 2.0, copyright 2025 FutureHouse. The code license covers the GitHub repository of reward functions and data utilities, not training code. The benchmark test set is tagged CC BY 4.0. The base model's card, published by Mistral AI's Hugging Face organization, also lists Apache 2.0.12456
Component reuse rights
- weights
- Reviewed qualifying license recorded — check scope and conditions
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Safetensors weights on Hugging Face without gating.21 |
| Inference codeIs code for running the model published? | Public | The README loads the model with Hugging Face Transformers; the card lists hosted inference providers and community quantizations for llama.cpp, Ollama, and LM Studio.31 |
| Training codeIs the code used to train the model published? | Unknown | The GitHub README states that the repository does not contain training code and points to third-party frameworks (NeMo-RL, TRL) for the SFT and RL phases.3 |
| Training-data informationDoes the information cover provenance, scope, acquisition, selection, labeling, processing, and where data or alternatives can be obtained? Access alone does not establish completeness. | Partial | The card lists the training task types, names EveBio as the source of receptor-binding data, and refers to the preprint for data details. This review confirms partial disclosure, not completeness.1 |
| Training-data accessCan the training data be obtained? This is independent of information completeness and reuse rights; original unshareable data need not be downloadable. | Unknown | Only the 325-question benchmark test set was found as a download; no training set was identified in the sources reviewed.53 |
| Complete training pipelineIs the complete base-training and preprocessing pipeline published, including configuration? Fine-tuning code or an inference SDK alone is insufficient. | Unknown | No complete training or preprocessing pipeline was found in the sources reviewed. |
| Legacy data assessment (v0.1)Historical assessment combining download access and disclosure. Preserved for traceability; excluded from the v0.2 tier calculation. See the new separate assessments above. | Unknown | Not assessed. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The card and README describe the stages (SFT on DeepSeek-R1 traces, GRPO specialists, rejection sampling, generalist SFT, all-task GRPO, safety post-training) at a high level; the preprint, which the card cites for details, was not reviewed for this record.13 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The benchmark test set is public, and the repository publishes the reward functions and an example script for scoring a model on it.35 |
What it is useful for
Chemistry tasks it was trained for, such as proposing molecules that meet a target pKa, solubility, scent, receptor-binding, or toxicity property, converting names or formulas to SMILES, one-step retrosynthesis, and reaction outcome prediction. The card says it is not a general-purpose chat model and cannot, for example, report the pKa of a given molecule.1
Run and use notes
Organization context
Provenance and derivatives
Derived from Mistral AI's Mistral-Small-24B-Instruct-2501. Per the model card, FutureHouse first fine-tuned the base model on reasoning traces from DeepSeek-R1, trained task-specific specialist models with GRPO and verifiable rewards, fine-tuned the base model again on filtered specialist reasoning traces, ran GRPO across all tasks, and then applied safety post-training. The base model is not a U.S.-developed model.126
- Derived from: Mistral-Small-24B-Instruct-2501 (Mistral AI) (external site: huggingface.co) — Base model named in the model card metadata and text; developed by Mistral AI.
Other releases in the ether0 family
No other releases in this family have been assessed.
What this catalog does not know
- Training code: unknown.
- Training-data access: unknown.
- Complete training pipeline: unknown.
- Legacy data assessment (v0.1): unknown.
Have a primary source? How to report a correction.
U.S. eligibility
Eligible · basis: U.S.-governed project
Published by FutureHouse in its Hugging Face organization; the card states a 2025 FutureHouse copyright. FutureHouse is a 501(c)(3) nonprofit in San Francisco per its about page and the IRS exempt-organization extract. The base model comes from Mistral AI and is recorded as provenance; it is not treated as a U.S.-developed model.1278
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.