Mochi 1 preview
Release in the Mochi family · version mochi-1-preview
A 10B-parameter text-to-video diffusion model built on Genmo's Asymmetric Diffusion Transformer (AsymmDiT) and released as a research preview in October 2024. It is paired with AsymmVAE, a 362M-parameter video autoencoder released alongside it, and encodes prompts with a single T5-XXL model. The initial release generates 480p video.214
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: genmoai/mochi repository (external site: github.com)
- Release notes: Mochi 1 announcement (external site: genmo.ai)
- Documentation: Mochi in Hugging Face Diffusers (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without gating; the repository also gives a magnet link and a download script.12
Availability is separate from permission: read the license before using or redistributing.
Genmo's announcement says the Apache 2.0 license permits personal and commercial use. The README notes that NSFW filtering is limited and advises organizations to add their own safety measures before deploying the weights in commercial services or products.42
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | The DiT and the AsymmVAE weights are both published.14 |
| Inference codeIs code for running the model published? | Public | Genmo's repository provides a Gradio UI, a CLI, and a Python API, and Hugging Face Diffusers includes a MochiPipeline.25 |
| Training codeIs the code used to train the model published? | Partial | The repository includes a LoRA fine-tuning trainer (added November 2024); this catalog found no release of the pretraining code.2 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Unknown | Genmo's announcement, README, and model card do not describe the training data.41 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Unknown | Not assessed. |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Unknown | Not assessed. |
What it is useful for
Text-to-video generation for research and development. Genmo notes that the model is optimized for photorealistic styles and does poorly on animated content, and that minor warping can occur with extreme motion. The repository includes a LoRA trainer for fine-tuning on one's own videos.2
Run and use notes
- Genmo's README says single-GPU use of its own implementation needs about 60GB of VRAM and recommends at least one H100. Per the Diffusers documentation, that implementation runs the text encoder and VAE at float32 precision and the DiT in BF16.25
- The Diffusers documentation says the full-precision weights need at least 42GB of VRAM and the bfloat16 variant 22GB (with a slight quality drop), both with model CPU offload and VAE tiling enabled; it also shows 8-bit loading with bitsandbytes and splitting the transformer across two 24GB GPUs.5
Organization context
Provenance and derivatives
Other releases in the Mochi family
No other releases in this family have been assessed.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.