Explainer
Choosing hosted access or local inference
Should I use an API or run a model on my own machine?
Reviewed Oct 1, 2026. General information, not legal or professional advice. All explainers
Hosted access sends your inputs over the network to a provider, which runs the model and returns the output. Local inference runs downloaded weights on hardware you control. Hosted access is quick to start and the provider maintains the hardware and software, but your inputs leave your machine and the provider's data policy governs them. Local inference keeps inputs on your hardware but needs a machine that can run the model and time to look after it. Hardware is paid for up front; hosted services charge for usage or by plan. Neither is cheaper in every case. This page is general information, not legal advice.
The same model, two ways
Some open-weight releases can be used either way. The gpt-oss-20b model card (external site: huggingface.co) shows how to run it locally with Ollama, and Ollama's library page (external site: ollama.com) also lists cloud versions of it. Ollama's cloud documentation (external site: docs.ollama.com) says cloud models run in Ollama's cloud and need no download. The weights are the same; what changes is where the computation happens, who can see the inputs, and whose terms apply.
Convenience and maintenance
Hosted access needs an account and usually an API key, and the provider runs and upgrades everything behind it. In exchange, the provider decides timing: it can retire a model. Ollama's cloud documentation, for example, notes upcoming cloud-model retirements and says downloaded local models are not affected.
Local inference means installing a runtime, downloading weights, keeping storage free, and applying updates. Ollama's FAQ (external site: docs.ollama.com) says its macOS and Windows apps download updates automatically while Linux users re-run the install script, and it lists each system's model folder. When something breaks, you troubleshoot it yourself.
Hardware cost and usage charges
Keep two kinds of cost apart. A one-time hardware cost is what you pay to own a machine that can run the model, plus power and upkeep. A usage charge is what a hosted service bills for requests, for the amount of text processed, or for a plan. Which is lower depends on how often you run the model, which model you need, what you already own, and the provider's prices on the day you compare. USASI does not publish prices; if you compare, use the provider's official pricing page and note the date.
Privacy boundaries
With hosted access, the boundary is the provider's data policy, so read its current version. Two examples, as published on October 1, 2026:
- OpenAI's data controls page (external site: developers.openai.com) says data sent to its API has not been used to train its models since March 1, 2023, unless the customer opts in. It also says abuse-monitoring logs, which may contain prompts and responses, are kept for up to 30 days by default, and that eligible customers can be approved for Zero Data Retention.
- Ollama's privacy policy (external site: ollama.com) says cloud prompts and responses are processed transiently and not used for training. For local use, it says Ollama does not collect, store, transmit, or have access to the content you process, though it may collect limited device and usage metadata such as the app version and request counts.
With local inference, the boundary is your machine and anyone who can reach it, so a runtime exposed on a network, or a shared computer, is not automatically private.
Network behavior of local tools
Local runtimes can still use the network. Check these points:
- Downloads. The llama.cpp server can fetch models from Hugging Face, and its
--offlineflag limits it to the local cache (server README (external site: github.com)). Hugging Face's libraries offerHF_HUB_OFFLINEfor the same purpose (environment variables (external site: huggingface.co)). - Telemetry. vLLM collects anonymous usage data by default, and its usage stats page (external site: docs.vllm.ai) explains how to opt out. Hugging Face libraries collect some usage data unless
HF_HUB_DISABLE_TELEMETRYorDO_NOT_TRACKis set. - Cloud features. Ollama offers local and cloud models in one app; its FAQ describes a local-only mode that turns off cloud models and web search.
- Listening addresses. Ollama binds to 127.0.0.1 on port 11434 by default, and the llama.cpp server listens on 127.0.0.1 with no API key unless one is set. Changing the address opens the model to other machines. vLLM's security guide (external site: docs.vllm.ai) warns that some components may listen on all network interfaces and that its
--api-keyoption protects only some endpoints, and it recommends a firewall.
Which terms apply
Locally, each component brings its own terms. For gpt-oss-20b, the weights come under the Apache License 2.0 (external site: huggingface.co) with a short usage policy (external site: huggingface.co) asking users to comply with applicable law. The runtime has a separate license: MIT for Ollama (external site: github.com) and llama.cpp (external site: github.com), Apache 2.0 for vLLM (external site: github.com). With hosted access, the provider's terms of service and usage policies also apply, even when the weights are openly licensed. OpenAI's data controls page, for example, cites its Usage Policies.
Worked example: a decision checklist
Jordan is a fictional reader invented for this page. Jordan coordinates volunteers at a small nonprofit, wants to summarize meeting notes that include volunteers' names a few times a month, and has a laptop but no IT staff.
- What goes in? Notes with personal names. Jordan first checks the organization's own rules. If a hosted service is allowed, Jordan reads that provider's current data policy, including how long logs are kept.
- Which model, under which terms? Jordan picks a release whose weights license permits this use, and reads the license file itself.
- Can the laptop run it? Jordan checks the publisher's stated memory needs and the runtime's supported hardware, then tests on the laptop.
- Who maintains it? Only Jordan, so a desktop app that updates itself fits better than a server setup.
- What touches the network? Jordan keeps the runtime on 127.0.0.1, turns on local-only mode where the runtime offers one, and checks telemetry settings.
- What will it cost? Use is occasional and the laptop is already paid for. If the model does not fit, Jordan would compare an upgrade with dated usage charges from official pricing pages.
Jordan's outcome: try local inference on the existing laptop first, and consider hosted access only after the organization approves a provider's data policy. Different answers, such as heavy daily use or no suitable hardware, could point the other way.
What you can do next
- Browse runtimes in the catalog, including Ollama, llama.cpp, vLLM, and MLX.
- Look up hosted API, self-hosting / local inference, and license scope in the glossary.
- Meet the people behind local AI.
- Read understanding inference hardware, how to read a model card, and what open weight and open source actually mean.
Sources
All read on October 1, 2026.
- OpenAI: Data controls in the OpenAI platform (external site: developers.openai.com); gpt-oss-20b model card (external site: huggingface.co), weights LICENSE (external site: huggingface.co), and USAGE_POLICY (external site: huggingface.co) on Hugging Face
- Ollama: FAQ (external site: docs.ollama.com), Cloud (external site: docs.ollama.com), Privacy Policy (external site: ollama.com) (last updated March 2026), gpt-oss library page (external site: ollama.com), LICENSE (external site: github.com)
- llama.cpp: server README (external site: github.com), LICENSE (external site: github.com)
- vLLM: Usage Stats Collection (external site: docs.vllm.ai), Security (external site: docs.vllm.ai), LICENSE (external site: github.com)
- Hugging Face: huggingface_hub environment variables (external site: huggingface.co)
Support Us
Help keep USASI useful.
Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project