Explainer
Understanding inference hardware
What determines whether a model can run on my hardware?
Reviewed Oct 1, 2026. General information, not legal or professional advice. All explainers
A model can run on your hardware when its weights, plus the working memory it needs while answering, fit in memory the runtime can use, and the runtime supports both your hardware and the model. Weight size depends on the parameter count and the precision used to store each parameter; quantization lowers the precision to save memory. Working memory grows with the context length and with the number of requests handled at once. Fitting is not the same as running well: a model split between GPU and system memory runs slower. Treat any single memory figure as a starting point, then test. This page is general information.
Which memory counts
A dedicated GPU has its own memory, called VRAM, separate from the computer's system RAM. When a model does not fit in VRAM, runtimes such as Ollama can split it between the two (FAQ (external site: docs.ollama.com)). Apple silicon works differently: the CPU and GPU share one pool. Apple's MLX framework is built around this unified memory model; its README (external site: github.com) says arrays live in shared memory and can be used on either device without transferring data. That is why hardware notes sometimes say "VRAM or unified memory"; in a shared pool, other running software competes for the same memory.
Model size, precision, and quantization
The weights take up roughly the number of parameters multiplied by the bits stored for each one. Granite 4.2 8B, for example, is published in bfloat16, a 16-bit format, according to its model card (external site: huggingface.co). Quantization stores weights in fewer bits. llama.cpp's README (external site: github.com) lists integer quantization from 1.5-bit to 8-bit for faster inference and lower memory use, and vLLM's conserving memory (external site: docs.vllm.ai) page notes the cost: lower precision.
For gpt-oss-20b, OpenAI's model card (external site: arxiv.org) says the mixture-of-experts weights, more than 90% of all parameters, were quantized to the MXFP4 format at 4.25 bits per parameter, and lists a 12.8 GiB checkpoint for 20.9B total parameters. Only about 3.6B parameters are active for each token, but the checkpoint holds all of them.
Context length and workload
A runtime keeps a cache for the text in the current context window, and that cache grows as the context gets longer. Ollama's context length page (external site: docs.ollama.com) says a larger context increases the memory required, and that its default context depends on available VRAM: 4k tokens below 24 GiB, 32k from 24 to 48 GiB, and 256k at 48 GiB or more. A model's documented maximum is not necessarily the default. Handling several requests at once multiplies the need; Ollama's FAQ says required memory scales with the number of parallel requests times the context length. Runtimes offer ways to trade quality or capacity for memory, such as Ollama's quantized cache types and vLLM's limits on context length and batch size.
Runtime support
Hardware and operating-system requirements differ by runtime:
- Ollama's hardware support page (external site: docs.ollama.com) lists NVIDIA GPUs with compute capability 5.0 or higher and driver 550 or newer, specific AMD GPUs through ROCm v7, Apple GPUs through Metal, and more GPUs through Vulkan.
- vLLM's GPU installation page (external site: docs.vllm.ai) requires Linux and, for NVIDIA, compute capability 7.5 or higher; it does not support Windows natively.
- llama.cpp's README lists CPU optimizations for Apple silicon, x86, and RISC-V, and backends including Metal, CUDA, HIP, Vulkan, and SYCL.
A GPU that one runtime supports may not work with another. Models add their own requirements: the gpt-oss-20b card (external site: huggingface.co) says the model works correctly only with OpenAI's harmony response format, and the Granite 4.2 8B card asks for vLLM v0.20 or later for its reasoning parser.
Why one memory number is not enough
A memory figure answers "will it load?" under particular assumptions, not "will it respond fast enough for my task?"
- Placement. OpenAI's guide to running gpt-oss with Ollama (external site: developers.openai.com) says you can offload to the CPU when short on VRAM but should expect slower responses. Ollama's context length page advises against offloading to the CPU when performance matters.
- Settings. llama.cpp's performance troubleshooting notes (external site: github.com) show one model on one machine generating text at markedly different speeds depending on thread count and GPU offload settings.
- Assumptions. A figure holds only for its stated precision, context length, and runtime.
- Your task. Long documents, coding agents, or several users need more context or more parallel requests than a short chat.
Worked example: a dated compatibility checklist
gpt-oss-20b with Ollama, checked October 1, 2026. USASI did not run the model, so nothing below is a USASI test. Each line is labeled Stated (by the publisher or runtime), Estimated (a method you apply), or Tested (what you measure).
- Release (Stated). openai/gpt-oss-20b on Hugging Face, revision 6cee5e8; in Ollama, the
gpt-oss:20btag. OpenAI's Ollama guide (August 5, 2025) says the models ship in MXFP4, with no other quantization offered. - Memory (Stated, not tested by USASI). OpenAI's model card, section 2.1: with the mixture-of-experts weights in MXFP4 (4.25 bits per parameter), the smaller model can "run on systems with as little as 16GB memory." Table 1 lists the checkpoint at 12.8 GiB. OpenAI's Ollama guide recommends at least 16GB of VRAM or unified memory.
- Download (Stated). Ollama's library page (external site: ollama.com) lists
gpt-oss:20bat 14GB with a 128K context window. That is the download size, not the memory needed while running. - Hardware support (Stated). Find your GPU and driver on Ollama's hardware support page. If it is not listed, do not assume GPU acceleration.
- Context (Stated). The model card says the context length was extended to 131,072 tokens. Ollama's default depends on your VRAM, and raising it increases memory use.
- Your budget (Estimated). Add the checkpoint size, room for the context length and parallel requests you plan to use, and what your system already uses. If the total exceeds your GPU memory, expect a split between GPU and CPU, or plan a shorter context.
- Your run (Tested). Run the model, then
ollama ps. Record the SIZE, the PROCESSOR column (100% GPU, 100% CPU, or a split), the CONTEXT, your Ollama version, and the date. Then try a typical prompt and judge whether the response time is acceptable. - Re-check. Repeat after updating the runtime or changing the model, context, or parallel settings.
What you can do next
- Browse runtimes, or only those with documented supported platforms.
- Read the run notes for gpt-oss-20b and Granite 4.2 8B.
- Look up quantization, context window, inference, and weights in the glossary.
- Meet the people behind local AI.
- Read choosing hosted access or local inference and how to read a model card.
Sources
All read on October 1, 2026.
- OpenAI: gpt-oss-120b and gpt-oss-20b Model Card (external site: arxiv.org) (August 5, 2025); gpt-oss-20b on Hugging Face (external site: huggingface.co) (revision 6cee5e8); How to run gpt-oss locally with Ollama (external site: developers.openai.com) (August 5, 2025)
- Ollama: Hardware support (external site: docs.ollama.com), Context length (external site: docs.ollama.com), FAQ (external site: docs.ollama.com), gpt-oss library page (external site: ollama.com)
- llama.cpp: README (external site: github.com), Token generation performance troubleshooting (external site: github.com)
- vLLM: GPU installation (external site: docs.vllm.ai), Conserving memory (external site: docs.vllm.ai)
- Apple: MLX README (external site: github.com)
- IBM: granite-4.2-8b model card (external site: huggingface.co)
Support Us
Help keep USASI useful.
Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project