SuperApriel-15B-Instruct
Release in the Apriel family · version SuperApriel-15B-Instruct
Maintained by ServiceNow (SLAM Labs)13
A 15B-parameter instruction-tuned "supernet" in which each of 48 decoder layers carries four token-mixer variants: full attention, sliding-window attention, Gated DeltaNet, and Kimi Delta Attention. Choosing one mixer per layer gives deployment presets that trade output quality for decoding speed from a single checkpoint. It was derived from Apriel-1.6-15b-Thinker by distillation followed by supervised fine-tuning.13
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Paper: Super Apriel: One Checkpoint, Many Speeds (arXiv 2604.19877) (external site: arxiv.org)
- Repository: Fast-LLM (training code and Apriel2 model implementation) (external site: github.com)
Availability and license
Overall availability
Downloadable from Hugging Face without gating. Loading uses custom model code shipped in the repository (trust_remote_code), and presets are selected by copying a preset configuration file.12
Availability is separate from permission: read the license before using or redistributing.
The weights license is stated as MIT in the model card's metadata and License section; the Hugging Face repository has no separate LICENSE file. Fast-LLM's LICENSE file applies Apache 2.0 to all files "unless otherwise noted".125
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | One supernet checkpoint with preset configurations is downloadable from Hugging Face.12 |
| Inference codeIs code for running the model published? | Public | The card documents inference with Hugging Face Transformers (custom modeling code in the repository) and with vLLM through a Fast-LLM plugin published on a feature branch.1 |
| Training codeIs the code used to train the model published? | Public | The paper states that the Fast-LLM training code is released; Fast-LLM's main branch contains an apriel2 model implementation.36 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The paper says distillation used the Apriel pretraining corpus and SFT data, and that the SFT stage used instruction-tuning data; the datasets are not documented as released.4 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The card and paper describe two-stage stochastic distillation from a frozen Apriel 1.6 teacher followed by targeted SFT, with token counts, batch shape, and GPU counts, and link training logs. This catalog did not confirm that complete configurations are published.14 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card and paper report benchmark results for each preset; the paper says it uses LM Evaluation Harness task formulations, but no evaluation configuration for this release was found.14 |
What it is useful for
The card lists code assistance, multi-step reasoning, question answering, function calling, instruction following, and agent use cases, and describes serving one preset for a fixed deployment or switching presets at runtime. It says the model is not intended for safety-critical use without human oversight.1
Run and use notes
- The model card recommends vLLM with the Fast-LLM plugin. It states that single-preset mode loads approximately 27 GiB of weights in bf16 and supernet mode (all four mixers per layer, allowing runtime preset switching) approximately 46 GiB in bf16, excluding KV cache.1
- Presets that use Gated DeltaNet or Kimi Delta Attention layers need the causal-conv1d and mamba-ssm packages when run with Transformers; attention-only presets do not.1
Organization context
Provenance and derivatives
Derived from ServiceNow's Apriel-1.6-15b-Thinker: shared parameters (feed-forward layers, embeddings, and norms) are inherited from Apriel 1.6, and the new mixer weights were trained by distillation from a frozen Apriel 1.6 teacher, then supervised fine-tuning. The Apriel 1.x line began from Mistral AI's Pixtral-12B-Base-2409, and the model keeps a Pixtral vision encoder.147
- Derived from: Apriel-1.6-15b-Thinker — Teacher model and source of the shared weights.
- Derived from: SuperApriel-15b-Base (external site: huggingface.co) — Stage 1 distilled supernet; no separate catalog record.
- Derived from: Pixtral-12B-Base-2409 (Mistral AI) (external site: huggingface.co) — Starting checkpoint of the Apriel 1.5 line.
Other releases in the Apriel family
- Apriel-1.6-15B-ThinkerModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S.-governed project
The model is published by ServiceNow's ServiceNow-AI organization on Hugging Face, and the accompanying report is credited to ServiceNow's SLAM Labs. ServiceNow, Inc. is a Delaware corporation headquartered in Santa Clara, California, per its 2025 Form 10-K. Its lineage traces through Apriel 1.6 and 1.5 to Pixtral-12B-Base-2409, published by Mistral AI; that base is recorded under provenance and is not treated as U.S.-developed.1387
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.