Treadstone Associates
Article · 7 min read

What open-weight AI models change

Some AI providers publish the trained parameters behind a model; most keep them private behind an API. What that difference actually changes for a business deciding between them.

Treadstone Associates · Updated 2026

Key takeaways

  • • An open-weight model is one whose trained parameters are published under a stated licence for anyone to download and run independently.
  • • “Open-weight” is narrower than “open source” — publishing weights doesn't guarantee the training data or training process is published too.
  • • Self-hosting changes where data goes and who controls versioning, and trades a vendor's operating burden for your own infrastructure cost.
  • • NIST's own risk framework treats a pre-trained model as a third-party component regardless of who hosts it — self-hosting does not remove the need to vet it.

What is actually being published, and what isn't

A model’s weights are the large set of internal numerical parameters that training produced — the thing described elsewhere in this hub as what makes a model able to predict text at all (see how large language models work). An “open-weight” model is one where the organization that trained it publishes those parameters for anyone to download and run on their own infrastructure, under a stated licence, rather than keeping them private and reachable only through a hosted API. Anthropic’s own product documentation shows the distinction in practice, inside a single table of embedding models it recommends: alongside several models offered only as a hosted service, one entry is labelled directly — (Anthropic, developer documentation — embeddings) “Open-weight model (Apache 2.0 license) available on Hugging Face” — sitting in the same list as models that are not open-weight at all. The label is doing real, specific work: it tells a developer that particular model can be downloaded and run independently, under a named licence, while its neighbours in the same table cannot.

This is a narrower claim than “open source” often implies in other software contexts. Publishing the weights lets someone run and adapt the trained model; it does not necessarily mean the training data, the training code, or the exact training process behind it is published alongside it. “Open-weight” and “fully open” are not guaranteed to be the same thing, and the licence attached to a specific release is what actually determines what you may do with it — which is why the licence, not the word “open”, is the detail worth reading.

What actually changes for a business that runs one

Running an open-weight model’s downloaded parameters on your own infrastructure, instead of calling a hosted API, changes a specific, identifiable set of things. Data no longer has to leave your own environment for every request, because inference happens wherever you deploy the weights — relevant wherever a business is weighing data residency or exposure, not merely convenience. You are not subject to a vendor changing or retiring the exact model version behind an API endpoint without notice, since you control the copy you run. And ongoing cost shifts from a per-call charge to the cost of the infrastructure required to run inference yourself — which, for a large model, is a real and specific engineering cost, not a free alternative to paying per request.

None of that is unambiguously “better” — it trades a vendor’s operational burden for your own, in exchange for more control. A hosted, closed model still generally receives ongoing updates and support directly from the provider without your infrastructure team doing anything; a self-hosted open-weight model does not update itself.

What does not change: the model is still a third-party component

Downloading a model’s weights does not make its provenance yours to vouch for. The U.S. National Institute of Standards and Technology names this directly in its generative-AI risk profile, under a risk category it calls Value Chain and Component Integration: (NIST AI 600-1 — United States) “GAI value chains involve many third-party components such as procured datasets, pre-trained models, and software libraries. These components might be improperly obtained or not properly vetted,” creating risk that travels with the component regardless of who is currently running it. Canada’s federal, provincial and territorial privacy commissioners make the same accountability point from the governance side, in the same joint principles that cite this NIST framework: an organization adopting any generative AI system, however it is deployed, is expected to (OPC — Principles for generative AI) “evaluate the validity and reliability of the generative AI tool for the intended purpose” before relying on it. Hosting a model yourself does not exempt you from that evaluation — if anything, it puts more of the vetting work on you, since a hosted vendor's own review is one thing you no longer have between you and the model.

The decision this actually maps onto

Choosing between an open-weight model you host and a closed model you call through an API is not really a decision about which is more capable in the abstract — capability varies release to release on both sides, and a genuine comparison is a benchmarking exercise this hub deliberately does not attempt. It is a decision about where you want operational responsibility to sit: with a vendor’s hosted service and its own update cycle, or with your own infrastructure and your own vetting of a third-party component you now run directly. That is a strategy and governance question before it is a technical one, which is why it is worth working through before committing either way — see our approach to AI strategy and roadmapping.

Why the licence attached to a specific release is the detail that actually matters

“Open-weight” is a description of distribution, not a single fixed set of permissions — two models both fairly described as open-weight can carry meaningfully different licences, and the licence is what actually determines what a business may do with a downloaded copy: whether it may be used commercially, whether outputs generated from it carry any obligation back to the original publisher, and whether a modified or fine-tuned version may be redistributed at all. The Apache 2.0 licence named on Anthropic’s own embeddings documentation for the open-weight model in its table is one specific, well-known set of terms; a different open-weight release elsewhere could carry a more restrictive licence limiting commercial use, and nothing about the phrase “open-weight” by itself tells you which you are getting. Reading the specific licence attached to a specific release, before building anything on top of it, is not a formality — it is the actual answer to “what am I allowed to do with this.”

Related, within this hub: how large language models work and what actually changes an AI’s output. If you are choosing between hosting your own model and using a vendor's API, see AI strategy and roadmapping.

Common questions

Is an open-weight model the same as open-source software?

Not necessarily. Open-weight means the trained parameters are published for anyone to download and run, under a stated licence. It does not guarantee the training data or training code is also published, which is a narrower claim than “open source” often carries elsewhere in software.

Does self-hosting an open-weight model remove third-party risk?

No. NIST's generative-AI risk profile treats a pre-trained model as a third-party component regardless of who is running it, and running it yourself shifts more of the vetting responsibility onto you rather than removing the need for it.

Is an open-weight model automatically cheaper than a hosted API?

Not automatically. It replaces a per-call charge with the cost of running your own infrastructure, which is a genuine engineering cost for a large model, not a free alternative.

Weighing whether to host a model yourself or use a vendor's API?

A short call is enough to map the trade-off against your own data and infrastructure.