Hub
Running AI on your own hardware
How running AI models on your own computer or device works: local runtimes, on-device frameworks, quantization, and what to check about an open-weight model before you run it.
Running a model locally means its weights are stored on a computer or device you control and the computation happens there. A runtime loads the model and makes it usable. llama.cpp is a C/C++ inference project with a command-line tool and an OpenAI-compatible server. Ollama runs open models and serves a REST API on the local machine. llamafile, now developed by Mozilla.ai, packages a model and a llama.cpp-based runtime into a single executable that runs without installation.124
Local does not automatically mean offline. Ollama also offers cloud models that run on Ollama's servers after you sign in, and its documentation says to disable cloud features if you want to use only local models. Its integrations can also connect a local setup to outside services such as messaging apps. Check a tool's settings before assuming that nothing leaves your machine.32
Hardware shapes what is practical. llama.cpp supports backends including CUDA for NVIDIA GPUs, HIP for AMD GPUs, Metal for Apple silicon, and Vulkan, and it can split a model between CPU and GPU when the model is larger than the GPU's memory. MLX, from Apple machine learning research, is an array framework for Apple silicon in which arrays live in memory shared by the CPU and GPU. ExecuTorch, PyTorch's on-device stack, captures a model with torch.export and runs it on phones, laptops, embedded systems, and microcontrollers.156
Quantization stores a model's weights at lower numerical precision to reduce memory use. llama.cpp supports integer quantization from 1.5 to 8 bits, and AWQ, from MIT's Han Lab, is a method for 3- and 4-bit weight-only quantization of large language models. Published results can depend on the precision used. OpenAI's gpt-oss-20b card says the model's mixture-of-experts weights were post-trained with MXFP4 quantization so that it runs within 16 GB of memory, and that all of its evaluations used that same quantization.178
Running a model on your own hardware does not change its license, so read the terms for the specific release before you rely on it. gpt-oss-20b, for example, is published under Apache 2.0. Its card describes it for lower-latency, local, or specialized uses and gives instructions for running it with Ollama or LM Studio on consumer hardware.8
What this hub covers
Covers running open-weight models on computers and devices you control: local runtimes and desktop apps, on-device frameworks, quantization methods, and models whose publishers document local use. It does not cover hosted APIs, hardware buying advice, or speed comparisons. Featured records are examples chosen to cover the subject, not a ranking or a complete list. Profiles of people in this area are on People Behind Local AI.
Reading path
- Hosted or local?Start here to decide whether running a model yourself fits what you need.
- Self-hostingA short definition of running software and models on infrastructure you control.
- What open weight and open source meanDownloadable weights are not the same as permission. Read this before choosing a model.
- QuantizationWhat lower-precision weights are and why they change memory needs.
- Inference hardwareHow memory, GPUs, and other accelerators affect which models are practical to run.
- How to read a model cardFind the license, intended uses, and any documented hardware notes for one release.
- People Behind Local AIProfiles of people whose public work helps others run models on their own hardware.
In the catalog
- Runtimes in the catalogSoftware that loads and runs models, including local runtimes and model servers.
- Records that mention local useA text search that matches records tagged for local inference.
- Open-weight model releasesReleases whose weights the public can obtain. Check each license before use.
Featured records
Examples chosen to cover the subject; not a ranking or a complete list.
- Element LabsOrganization
- Mozilla.aiOrganization
- llama.cppOpen model or tool
- OllamaOpen model or tool
- llamafileOpen model or tool
- MLXOpen model or tool
- ExecuTorchOpen model or tool
- AWQ (Activation-aware Weight Quantization)Open model or tool
- gpt-oss-20bOpen model or tool
- Georgi GerganovPeople Behind Local AI
- Awni HannunPeople Behind Local AI
- Tim DettmersPeople Behind Local AI
Primary documents
- llama.cpp README (external site: raw.githubusercontent.com)ggml-org (GitHub)The project's own list of hardware backends and quantization formats. Use it to check whether your machine is supported before downloading a model.
- Cloud (Ollama documentation) (external site: docs.ollama.com)OllamaExplains which Ollama models run on Ollama's servers instead of your machine, and how to keep to local models only.
- llamafile README (external site: raw.githubusercontent.com)Mozilla.ai (GitHub)Describes the single-file packaging approach and its basis in llama.cpp and Cosmopolitan Libc.
- MLX README (external site: raw.githubusercontent.com)Apple, ml-explore (GitHub)Apple's description of MLX and its unified memory model, which matters for how models use memory on Apple silicon.
- ExecuTorch README (external site: raw.githubusercontent.com)PyTorch (GitHub)Shows the export-then-run workflow for phones and embedded devices, and warns that accuracy, memory use, and performance must be validated on the target device.
- AWQ: Activation-aware Weight Quantization (README) (external site: raw.githubusercontent.com)MIT Han Lab (GitHub)The authors' own description of the AWQ weight-quantization method, with the code and 4-bit kernels that implement it.
- gpt-oss-20b model card (external site: huggingface.co)OpenAI (Hugging Face)A model card that ties a memory figure to a stated quantization and names local runtimes, a useful example of what to look for in a card.
Sources · reviewed Oct 1, 2026
Support Us
Help keep USASI useful.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project