Independent project. Not a U.S. government website.

USASI

Hub

Running AI on your own hardware

How running AI models on your own computer or device works: local runtimes, on-device frameworks, quantization, and what to check about an open-weight model before you run it.

Running a model locally means its weights are stored on a computer or device you control and the computation happens there. A runtime loads the model and makes it usable. llama.cpp is a C/C++ inference project with a command-line tool and an OpenAI-compatible server. Ollama runs open models and serves a REST API on the local machine. llamafile, now developed by Mozilla.ai, packages a model and a llama.cpp-based runtime into a single executable that runs without installation.124

Local does not automatically mean offline. Ollama also offers cloud models that run on Ollama's servers after you sign in, and its documentation says to disable cloud features if you want to use only local models. Its integrations can also connect a local setup to outside services such as messaging apps. Check a tool's settings before assuming that nothing leaves your machine.32

Hardware shapes what is practical. llama.cpp supports backends including CUDA for NVIDIA GPUs, HIP for AMD GPUs, Metal for Apple silicon, and Vulkan, and it can split a model between CPU and GPU when the model is larger than the GPU's memory. MLX, from Apple machine learning research, is an array framework for Apple silicon in which arrays live in memory shared by the CPU and GPU. ExecuTorch, PyTorch's on-device stack, captures a model with torch.export and runs it on phones, laptops, embedded systems, and microcontrollers.156

Quantization stores a model's weights at lower numerical precision to reduce memory use. llama.cpp supports integer quantization from 1.5 to 8 bits, and AWQ, from MIT's Han Lab, is a method for 3- and 4-bit weight-only quantization of large language models. Published results can depend on the precision used. OpenAI's gpt-oss-20b card says the model's mixture-of-experts weights were post-trained with MXFP4 quantization so that it runs within 16 GB of memory, and that all of its evaluations used that same quantization.178

Running a model on your own hardware does not change its license, so read the terms for the specific release before you rely on it. gpt-oss-20b, for example, is published under Apache 2.0. Its card describes it for lower-latency, local, or specialized uses and gives instructions for running it with Ollama or LM Studio on consumer hardware.8

What this hub covers

Covers running open-weight models on computers and devices you control: local runtimes and desktop apps, on-device frameworks, quantization methods, and models whose publishers document local use. It does not cover hosted APIs, hardware buying advice, or speed comparisons. Featured records are examples chosen to cover the subject, not a ranking or a complete list. Profiles of people in this area are on People Behind Local AI.

Reading path

  1. 1.Hosted or local?Start here to decide whether running a model yourself fits what you need.
  2. 2.Self-hostingA short definition of running software and models on infrastructure you control.
  3. 3.What open weight and open source meanDownloadable weights are not the same as permission. Read this before choosing a model.
  4. 4.QuantizationWhat lower-precision weights are and why they change memory needs.
  5. 5.Inference hardwareHow memory, GPUs, and other accelerators affect which models are practical to run.
  6. 6.How to read a model cardFind the license, intended uses, and any documented hardware notes for one release.
  7. 7.People Behind Local AIProfiles of people whose public work helps others run models on their own hardware.

In the catalog

Examples chosen to cover the subject; not a ranking or a complete list.

Primary documents

Sources · reviewed Oct 1, 2026

  1. 1.
    llama.cpp README (external site: raw.githubusercontent.com)

    ggml-org (GitHub) · Repository · accessed Oct 1, 2026

  2. 2.
    Ollama README (external site: raw.githubusercontent.com)

    Ollama (GitHub) · Repository · accessed Oct 1, 2026

  3. 3.
    Cloud (Ollama documentation) (external site: docs.ollama.com)

    Ollama · Documentation · accessed Oct 1, 2026

  4. 4.
    llamafile README (external site: raw.githubusercontent.com)

    Mozilla.ai (GitHub) · Repository · accessed Oct 1, 2026

  5. 5.
    MLX README (external site: raw.githubusercontent.com)

    Apple, ml-explore (GitHub) · Repository · accessed Oct 1, 2026

  6. 6.
    ExecuTorch README (external site: raw.githubusercontent.com)

    PyTorch (GitHub) · Repository · accessed Oct 1, 2026

  7. 7.
  8. 8.
    gpt-oss-20b model card (external site: huggingface.co)

    OpenAI (Hugging Face) · Model card · accessed Oct 1, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project