Search

· Strategy · 12 min read

Open Weights, Frontier APIs or Self-Hosting: A Workload-Based Decision Matrix

Open weights vs frontier APIs vs self-hosting: a decision matrix built on request volume, data sensitivity, eval maturity and compliance.

Featured image for: Open Weights, Frontier APIs or Self-Hosting: A Workload-Based Decision Matrix

In my work with teams choosing how to run language models, the argument almost always starts in the wrong place. Someone asks whether open models are “good enough yet”, or whether the frontier labs are “too expensive”, and the room splits into camps. Nobody asks what the workload actually looks like. I’ve watched a team spend a quarter standing up GPUs for a feature that handled a few thousand requests a day, and I’ve watched another pay a steady and growing API bill for a classification job that a small model on one box could have done.

Both mistakes have the same root. The choice between a frontier API, an open-weights model served by someone else, and a stack you run yourself is not a question about which technology is better. It is a question about four properties of your situation: how many requests you send, how sensitive the data is, how well your team can measure model quality, and what your compliance posture requires. Change those and the right answer changes with them.

There are really three options here, and they are easy to blur together. A frontier API means calling a closed lab’s hosted model. Open weights through a provider means using a model whose weights are published, but running it on someone else’s managed endpoint. Self-hosting means you run the weights on hardware you control, whether owned or rented. The second and third share a model family and differ in who carries the operational burden.

Key Takeaways
  • Pick by workload, not by camp: request volume, data sensitivity, evaluation maturity and compliance posture decide the answer, and each option wins somewhere.
  • Frontier APIs fit low-to-moderate volume, hard reasoning and teams that need vendor support and built-in safety tooling; the cost is less control and a per-token bill that scales with usage.
  • Self-hosting fits steady high volume, strict data residency and teams that can run their own evaluation and operations; the cost is staffing, utilization risk and owning the safety layer.
  • Managed open weights is the middle path, and the thresholds in this article are rules of thumb to calibrate against your own numbers, not benchmarks.

The four variables that matter

Before the matrix, the inputs. I’ll give rough thresholds throughout. Treat them as starting points I use for a first conversation, not measured industry constants. Prices, hardware and model quality shift often enough that you should redo the arithmetic with current quotes before committing.

Request volume

Volume decides whether fixed costs can be spread thin enough to matter. An API charges per token and has no floor. Self-hosting has a floor: hardware or reserved capacity, plus the people to run it, whether or not a single request arrives. So the question is not “which is cheaper per token” but “at what utilization does my fixed cost beat my variable cost”.

My rough heuristic: below a few million requests a month, the engineering time to run a serving stack almost always costs more than the API bill it saves. Between that and tens of millions, managed open weights often wins because you get cheaper tokens without owning the cluster. Beyond that, and only if traffic is steady rather than spiky, self-hosting starts to pay, because you can keep accelerators busy. The word “steady” matters. A cluster sized for peak and idle at night is an expensive way to lose an argument with your finance team.

Data sensitivity

Ask what happens if a prompt leaks, is retained, or crosses a border. For public or lightly sensitive content, a contractual no-training, limited-retention API agreement from a reputable provider is often enough. For regulated records, trade secrets, or anything where the legal answer to “where did this data go” must be “nowhere outside our boundary”, the set of acceptable options narrows quickly. Self-hosting inside your own network, or a managed open-weights deployment inside your cloud account, answers that cleanly. A third-party API may still be permissible, and regulated data can sometimes go to a frontier vendor under a suitable agreement or a private deployment inside your own cloud account, but you will spend real time on the paperwork. The flow above reflects this: a vendor path stays open when your legal and compliance teams accept the terms.

Evaluation maturity

This is the variable people skip, and it is the one I would weight most. Can your team tell, with a repeatable test set and a scoring method, whether a model change made things better or worse? If yes, you can swap models, tune smaller ones, and trust your cost savings. If no, every migration is a guess dressed as engineering.

Frontier APIs are forgiving here because the strongest models tolerate sloppy prompts and fuzzy requirements. Open and self-hosted models usually reward you for investing in evaluation, because you will be choosing among many candidates, quantizations and fine-tunes. A team with no eval harness that self-hosts is taking on the hardest part of the job with the fewest tools.

Compliance posture

Compliance is partly about data location and partly about proof. Auditors want to know what model ran, which version, what it saw and who could access it. Pinned open weights you host give you a stable, inspectable artifact. A hosted frontier model can change underneath you on the provider’s schedule, though providers typically offer version pinning and deprecation notices, and many hold certifications that save you months of audit work. Neither is automatically safer. The question is which one your auditor and your sector regulator will accept with the least friction.

The decision flow

Here is the sequence I walk through. It is a first pass, and every branch has exceptions.

Supplemental Explainer

Notice that the flow never ends at “use the best model” by default. Every path to self-hosting or managed open weights passes through a capability check: if open models cannot clear your quality bar on your own evals, the flow ends at a frontier API or a private frontier deployment, and that is a permanent, legitimate outcome rather than a stopgap. The only stopgap is the frontier API you use while building evals. Most workloads today do not need the most capable model available, and a few genuinely do.

Where each option is the better fit

Frontier APIs

Best fit: low-to-moderate volume, tasks at the edge of what models can do, and teams that want to ship this month. Think complex multi-step reasoning, hard coding tasks, long-document analysis, or a new product where you do not yet know what quality you need. You get the highest capability ceiling, no infrastructure, and usually mature safety tooling: content filters, abuse monitoring, red-team investment and documented behavior on risky requests. You also get a vendor to call, with service levels, when something breaks at two in the morning.

Counterweight: you give up control. Pricing, rate limits, model versions and retention terms are the provider’s to set within the contract. Data leaves your boundary, and some regulated workloads cannot accept that. None of this makes the option bad. It makes it the right tool when speed and capability matter more than control, and the wrong one when volume is high and the task is simple.

Managed open weights

Best fit: the middle of the volume range, workloads that do not need the very top of the capability curve, and teams that want portability. You can run a published model through a provider or inside your own cloud account, keep the option to move, and often pay noticeably less per token than a top-tier API for tasks like extraction, summarization, routing and classification. If your data has to stay in a particular region or account, a managed deployment inside your tenancy can satisfy that without you operating GPUs.

Counterweight: you inherit the safety and quality work. Open models ship with varying levels of alignment effort and documented safeguards, and the guardrails around them are yours to assemble: input and output filtering, abuse handling, prompt injection defenses. Support is also thinner. The provider supports the endpoint, not the behavior of the model, so when quality drifts for your use case, you are the one with the test set that proves it.

Self-hosting

Best fit: steady high volume, strict data residency or air-gapped environments, and deep customization such as fine-tuning on proprietary data or tight latency control. A team that already runs GPU workloads, has an evaluation habit, and sees predictable traffic can get the lowest marginal cost and the most control over versions, behavior and data flow. This is also the cleanest answer when the requirement is literally “no data leaves the building”.

Counterweight: you own everything. Capacity planning, upgrades, kernel and driver issues, batching and caching, observability, security patching, and the safety layer. Utilization risk is the quiet killer. If demand falls or the model you built around is superseded, you hold the hardware commitment. And the people who can do this well are scarce, so count their salaries, not just the rack, in your comparison. Self-hosting is a strong choice for the right team and an expensive hobby for the wrong one.

A worked example

Take a mid-sized insurer with three workloads, which is a composite of patterns I have seen rather than one specific client.

The first is claims triage: classifying inbound documents into a dozen categories and extracting a few fields. Volume is high and steady, the documents contain personal data, and the task is simple. Data sensitivity and volume both point away from a public API. A small open-weights model, fine-tuned and self-hosted or run in their own cloud account, fits. They already have a labeled history, so building the eval set is cheap.

The second is an internal assistant for underwriters who need to reason across long, messy policy documents. Volume is modest, a few hundred users, and quality is the entire point. A frontier API with a strong enterprise agreement fits better, provided the legal team signs off on the data terms. Running a comparable stack themselves would cost more in engineers than the API costs in tokens.

The third is customer-facing chat for simple policy questions. It is mid-volume, spiky, and carries brand risk if the bot says something wrong. Here managed open weights behind their own guardrails is reasonable, but a frontier model with its built-in safety tooling is a defensible alternative, and the decision depends on whether their team is ready to own the filtering layer. If the answer is “not yet”, start with the API and revisit.

Three workloads, three different answers, one company. That is the usual outcome, and it is why a single organization-wide policy of “we are an open-weights shop” or “we only use the frontier” tends to age badly.

What this means for leaders

If you sit above the engineering team, three questions cut through most of the noise.

First, what is the monthly cost curve if this feature succeeds tenfold? A flat curve favors fixed-cost options, and a curve you can tolerate favors the API. Second, can your team show you, with numbers, that a cheaper model is good enough for this task? If they cannot, the savings are speculative. Third, what does the auditor or regulator need to see, and which option produces that evidence with the least effort?

Procurement comfort exists on both sides: open-weights deployments offer inspectable artifacts and portability, and frontier labs have invested heavily in safety evaluation, certifications and enterprise support, which is real value for teams who would otherwise build those controls themselves. Neither changes the arithmetic for your workload.

A practical way to start

Pick your two highest-volume workloads and your one highest-risk workload. For each, write down the four variables with real numbers: monthly requests, data classification, whether an eval set exists, and the compliance constraint. Then run the same test set against a frontier model and one or two open candidates.

Revisit the decision every few quarters. Model quality, hardware cost and provider terms all move, and a decision that was right for a workload last year may not be now, in either direction. The skill worth building is not loyalty to a camp. It is the habit of measuring your own workload and letting the numbers pick the tool. This post obviously doesn’t cover ideas around model alignment which is yet another complex topic and assumes you undertstand that AI is your new security permeter. Leave your questions here on the topic and I would be happy to cover it in a lot more detail.

Enjoying this insight?

Join the distribution list to get deep dives on AI transitions and agency economics directly in your inbox. No spam, ever.

Back to Blog

Related Posts

View All Posts »
Distillation in 2026: The Hidden Arms Race

Distillation in 2026: The Hidden Arms Race

All frontier labs have converged on distillation as their primary post-training method. The real innovation is no longer in pre-training but in post-training pipeline design. Here is why distillation...

Governance: The "Human in the Loop" Fallacy

Governance: The "Human in the Loop" Fallacy

Humans cannot keep pace with AI outputs at scale. Here is why enterprise growth relies heavily on Constitutional AI, rather than just throwing more human reviewers at the problem.