The Honest TCO: Self-Hosting vs. API (with interactive calculator)
When does your own AI hardware actually pay back? The honest total-cost-of-ownership comparison with an interactive payback calculator: hardware, power, cooling, ops, redundancy vs. the API meter.
Note: Every figure in this article and in the calculator is an illustrative, order-of-magnitude model placeholder, not an offer and not a measured client result. Each default is anchored to a public market benchmark (published open-weight GPU hardware prices, German industrial electricity rates around 0.20 EUR/kWh, the standard three-year German IT depreciation convention, typical fully-loaded MLOps staffing cost, published per-token API pricing). Put in your own current numbers.
The decision in one sentence
Most “self-host vs. cloud” calculators compare the hardware price against monthly API spend and declare victory in six months. That comparison is wrong because it leaves out everything it takes to turn a box under a desk into a service you can rely on.
An honest TCO has two sides. This article counts both in full and gives you a model you can fill in yourself.
Part 1: The model (the honest part)
Side A: Self-hosting, full run-rate
| Cost component | What it really is | Why it is often forgotten |
|---|---|---|
| Hardware (depreciation) | The GPU node(s), spread over their useful life (German IT depreciation: typically around 3 years). | Booked as a one-off, so it looks “free” after purchase. It is not: it ages out. |
| Power | The rig draws watts around the clock if you want it available. | Invisible until the electricity bill arrives. |
| Cooling | Every watt of compute becomes heat that has to leave the room. | Folded into “the building” and never attributed. |
| Ops time | Patching, monitoring, model updates, incident response: a fraction of a real person. | The single biggest hidden cost, and the one that decides the outcome. |
| Redundancy | What you spend to not be down when one box fails. | The gap between “a nice demo” and “production”. |
| Setup (one-off) | Initial integration, security, evaluation per workload. | Real engineering effort before anything runs in production. |
Side B: The API meter
| Cost component | What it really is |
|---|---|
| Usage | Tokens in plus tokens out times the provider rate, summed over the month. |
| Growth | The bill scales with adoption: every new use case adds to it, quietly. |
| Floor | Near-zero fixed cost. You pay for what you use, including 0 EUR in a quiet month. |
The payback formula
One-off cost to stand up self-hosting (hardware + setup)
Payback (months) = ───────────────────────────────────────────────────────────
Monthly API spend − Monthly self-host run-rate (opex)
Two honest consequences fall straight out of this formula:
- If the monthly run-rate is close to the API spend, payback is never. Ops time alone can eat most of a small workload’s “savings”. For one light workload, self-hosting often does not pay back.
- Payback collapses as you add workloads. Power, cooling, redundancy and (mostly) ops time are largely fixed per node. Put a second and third workload on the same hardware and the API side grows while the self-host side barely moves. Self-hosting is a consolidation play, not a single-app play.
That is the real thesis, and it is the honest one: self-hosting wins when you have enough AI work to keep one well-run node busy.
Part 2: Run the numbers yourself
All defaults below are the illustrative figures from the worked example and are clearly marked editable. Put in your own numbers and see where your tipping point lands. The faint dashed line shows the “hardware-only” math, deliberately there as the contrast to the honest line.
Self-hosting (own hardware)
API (cloud LLM)
The illustrative worked example
Setup (illustrative): a desk-scale node in the spirit of what we run, roughly 20,000 EUR of hardware (around 4 DGX-Spark-class boxes), suitable for a mid-size open-weight model.
Scenario 1: One workload (the cautionary case)
| Item | Illustrative monthly | Basis |
|---|---|---|
| Hardware depreciation | 555 EUR | 20,000 EUR over 36 months |
| Power | 175 EUR | around 1.2 kW average, 24/7, 0.20 EUR/kWh |
| Cooling | 55 EUR | around 30 % overhead on IT power |
| Ops time | 1,100 EUR | around 0.15 FTE of an MLOps/admin, loaded |
| Redundancy (amortised) | 150 EUR | spare-capacity / HA provision |
| Self-host run-rate | around 2,035 EUR / mo | |
| Comparable API spend | around 3,000 EUR / mo | representative mid-size workload |
| Monthly difference | around 965 EUR / mo | |
| One-off setup (integration, security, eval) | around 10,000 EUR |
Payback around 30,000 EUR / 965 EUR per month, so around 31 months. Not the six-to-seven months the naive “hardware vs. API bill” math promises. With full-TCO honesty, payback for a single light workload is slow, well over two years, and may not be worth it on cost alone. (Data residency and independence may still justify it.)
The widely quoted “around 6 to 7 month” payback (around 20,000 EUR hardware / around 3,000 EUR per month) is the hardware-only view. We show it side by side precisely to make the gap between the easy story and the honest one visible.
Scenario 2: Three workloads on the same node (the real case)
The ops, power, cooling and redundancy are largely shared; only modest setup is added per workload.
| Item | Illustrative monthly |
|---|---|
| Self-host run-rate (shared node, 3 workloads) | around 2,400 EUR / mo |
| Comparable API spend (3x around 3,000 EUR) | around 9,000 EUR / mo |
| Monthly difference | around 6,600 EUR / mo |
| One-off setup (hardware + 3x integration) | around 40,000 EUR |
Payback around 40,000 EUR / 6,600 EUR per month, so around 6 months. Same hardware, same operating discipline, but now the consolidation does the work. This is where self-hosting earns its place.
Part 3: The honest footnote
We run this ourselves, so we will tell you the part the calculator cannot: operating it is the project. Opteria runs exactly one node with one GPU plus a staging environment: small, but real. We do not sell borrowed scale; we sell the method and the outcome. The reason our number for “ops time” is not zero is that we pay it every month, and we would rather you budget for it honestly than discover it after the hardware arrives.
So use the calculator for the cost story. Then ask the harder questions it does not price: Do you have enough AI workload to fill a node? Who owns uptime at 2 a.m.? How sensitive is the data? Those decide self-host vs. cloud at least as much as the euros do, and they are exactly the per-workload decision we make at Opteria.
Related: our broader Cloud-LLM vs. own infrastructure cost calculation and the decision tree: self-host or cloud.
Ready to implement AI in production?
We analyse your process and show you in 30 minutes which workflow delivers the highest ROI.