AWQ (Activation-aware Weight Quantization)
Project record
Maintained by MIT HAN Lab126
AWQ is a post-training method and codebase for low-bit (INT3/INT4) weight-only quantization of large language models, including instruction-tuned and multimodal models. It uses activation statistics to find the most important weight channels and scales them before quantization, without backpropagation. The repository also includes 4-bit CUDA kernels and TinyChat, an inference interface for running quantized models on desktop and edge GPUs.51
- Repository: Repository (mit-han-lab/llm-awq) (external site: github.com)
- Website: Project page (external site: hanlab.mit.edu)
- Paper: Paper: AWQ (arXiv 2306.00978) (external site: arxiv.org)
- License: LICENSE (MIT) (external site: github.com)
- Dataset hub: AWQ model zoo (pre-computed search results) (external site: huggingface.co)
Availability and license
Overall availability
Source code is public on GitHub and installed from source; pre-computed AWQ search results are published as a Hugging Face dataset.1
Availability is separate from permission: read the license before using or redistributing.
Component reuse rights
- weights
- Unknown — no complete fact-level rights review
- code
- Reviewed qualifying license recorded — check scope and conditions
- data
- Unknown — no complete fact-level rights review
- documentation
- Unknown — no complete fact-level rights review
No complete system-rights review is recorded for this release.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Public GitHub repository under the MIT License.12 |
| DocumentationIs user documentation published? | Public | The README covers installation, the model zoo, usage of the AWQ search and evaluation commands, and example scripts; the method is described in the MLSys 2024 paper.15 |
| InstallationAre installation instructions or packages publicly available? | Public | Installed from source: clone the repository, create a Python 3.10 conda environment, pip-install the package, build the CUDA kernels, and install FlashAttention. Separate steps are documented for NVIDIA Jetson Orin.1 |
| Supported platformsAre supported operating systems or hardware documented? | Public | The README targets NVIDIA GPUs through CUDA kernels, with examples on desktop GPUs (RTX 4090), Jetson Orin edge devices, and NVIDIA laptops. The README documents no other hardware backends.1 |
| Release statusAre versioned releases published? | Not public | The GitHub releases page lists no releases; the code is used from the main branch. The README's news entries run through April 2025.31 |
What it is useful for
Quantizing language and vision-language models to 4-bit or 3-bit weights to reduce the memory needed to serve them, and running the quantized models with the included kernels and TinyChat. The README provides example scripts for model families such as Llama, Qwen, DeepSeek-R1-Distill, and VILA, and pre-computed AWQ search results on Hugging Face.1
Run and use notes
- The README documents installation with a Python 3.10 conda environment, compiling the W4A16 CUDA kernels, and installing FlashAttention. For Jetson Orin it says to install NVIDIA's precompiled PyTorch (2.0.0 or later) and adjust the Python version to the JetPack release.1
- The README states that AWQ has been integrated into Hugging Face Transformers, vLLM, TensorRT-LLM, Hugging Face TGI, LMDeploy, and other tools, and points to AutoAWQ as a separate third-party implementation.1
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The code is published and maintained in the MIT HAN Lab's mit-han-lab GitHub organization, and the repository's MIT License names MIT HAN Lab as copyright holder. The lab is part of the Massachusetts Institute of Technology, which the IRS lists as a 501(c)(3) school in Cambridge, Massachusetts. The AWQ project page also credits researchers from Tsinghua University and the MIT-IBM Watson AI Lab; eligibility rests on the documented maintaining lab, not on authorship.12647
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.