Independent project. Not a U.S. government website.

USASI
SoftwareFramework

AWQ (Activation-aware Weight Quantization)

Project record

Maintained by MIT HAN Lab126

AWQ is a post-training method and codebase for low-bit (INT3/INT4) weight-only quantization of large language models, including instruction-tuned and multimodal models. It uses activation statistics to find the most important weight channels and scales them before quantization, without backpropagation. The repository also includes 4-bit CUDA kernels and TinyChat, an inference interface for running quantized models on desktop and edge GPUs.51Fact reviewed Oct 1, 2026

Last reviewedEntry updated Documented release Unknown

Availability and license

Overall availability

Public

Source code is public on GitHub and installed from source; pre-computed AWQ search results are published as a Hugging Face dataset.1

Availability fact review: Oct 1, 2026

Availability is separate from permission: read the license before using or redistributing.

The MIT License covers this repository's code. It does not cover the model weights that users quantize with AWQ.21Fact reviewed Oct 1, 2026

Component reuse rights

A readable or downloadable component is not automatically reusable. These indicators concern recorded license evidence, not system certification.
weights
Unknown — no complete fact-level rights review
code
Reviewed qualifying license recorded — check scope and conditions
data
Unknown — no complete fact-level rights review
documentation
Unknown — no complete fact-level rights review

No complete system-rights review is recorded for this release.

Public materials checklist

Items for a framework under USASI rubric v0.2. Unknown means unassessed or insufficient evidence.
Public materials checklist for AWQ (Activation-aware Weight Quantization)
ItemStatusNotes and evidence
Source codeIs the source code publicly readable?PublicPublic GitHub repository under the MIT License.12
DocumentationIs user documentation published?PublicThe README covers installation, the model zoo, usage of the AWQ search and evaluation commands, and example scripts; the method is described in the MLSys 2024 paper.15
InstallationAre installation instructions or packages publicly available?PublicInstalled from source: clone the repository, create a Python 3.10 conda environment, pip-install the package, build the CUDA kernels, and install FlashAttention. Separate steps are documented for NVIDIA Jetson Orin.1
Supported platformsAre supported operating systems or hardware documented?PublicThe README targets NVIDIA GPUs through CUDA kernels, with examples on desktop GPUs (RTX 4090), Jetson Orin edge devices, and NVIDIA laptops. The README documents no other hardware backends.1
Release statusAre versioned releases published?Not publicThe GitHub releases page lists no releases; the code is used from the main branch. The README's news entries run through April 2025.31

What it is useful for

Quantizing language and vision-language models to 4-bit or 3-bit weights to reduce the memory needed to serve them, and running the quantized models with the included kernels and TinyChat. The README provides example scripts for model families such as Llama, Qwen, DeepSeek-R1-Distill, and VILA, and pre-computed AWQ search results on Hugging Face.1Fact reviewed Oct 1, 2026

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README documents installation with a Python 3.10 conda environment, compiling the W4A16 CUDA kernels, and installing FlashAttention. For Jetson Orin it says to install NVIDIA's precompiled PyTorch (2.0.0 or later) and adjust the Python version to the JetPack release.1
  • The README states that AWQ has been integrated into Hugging Face Transformers, vLLM, TensorRT-LLM, Hugging Face TGI, LMDeploy, and other tools, and points to AutoAWQ as a separate third-party implementation.1

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The code is published and maintained in the MIT HAN Lab's mit-han-lab GitHub organization, and the repository's MIT License names MIT HAN Lab as copyright holder. The lab is part of the Massachusetts Institute of Technology, which the IRS lists as a 501(c)(3) school in Cambridge, Massachusetts. The AWQ project page also credits researchers from Tsinghua University and the MIT-IBM Watson AI Lab; eligibility rests on the documented maintaining lab, not on authorship.12647

Assessed Oct 1, 2026

Sources

  1. 1.
    mit-han-lab/llm-awq (GitHub repository and README) (external site: github.com)

    MIT HAN Lab · Repository · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  2. 2.
    llm-awq LICENSE (MIT License) (external site: raw.githubusercontent.com)

    MIT HAN Lab · License · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  3. 3.
    Releases · mit-han-lab/llm-awq (external site: github.com)

    MIT HAN Lab · Release notes · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  4. 4.
    AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (project page) (external site: hanlab.mit.edu)

    MIT HAN Lab · Official page · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  5. 5.
    AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv 2306.00978) (external site: arxiv.org)

    arXiv (MIT HAN Lab and co-authors) · Paper · published Jun 1, 2023 · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  6. 6.
    MIT HAN Lab (external site: hanlab.mit.edu)

    MIT HAN Lab · Official page · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

  7. 7.
    Exempt Organizations Business Master File Extract - Massachusetts (eo_ma.csv) (external site: irs.gov)

    Internal Revenue Service · Filing · accessed Oct 1, 2026 · evidence reviewed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project