The Honest TCO: Self-Hosting vs. API (with interactive calculator)

When does your own AI hardware actually pay back? The honest total-cost-of-ownership comparison with an interactive payback calculator: hardware, power, cooling, ops, redundancy vs. the API meter.

Note: Every figure in this article and in the calculator is an illustrative, order-of-magnitude model placeholder, not an offer and not a measured client result. Each default is anchored to a public market benchmark (published open-weight GPU hardware prices, German industrial electricity rates around 0.20 EUR/kWh, the standard three-year German IT depreciation convention, typical fully-loaded MLOps staffing cost, published per-token API pricing). Put in your own current numbers.

The decision in one sentence

Most “self-host vs. cloud” calculators compare the hardware price against monthly API spend and declare victory in six months. That comparison is wrong because it leaves out everything it takes to turn a box under a desk into a service you can rely on.

An honest TCO has two sides. This article counts both in full and gives you a model you can fill in yourself.

Part 1: The model (the honest part)

Side A: Self-hosting, full run-rate

Cost componentWhat it really isWhy it is often forgotten
Hardware (depreciation)The GPU node(s), spread over their useful life (German IT depreciation: typically around 3 years).Booked as a one-off, so it looks “free” after purchase. It is not: it ages out.
PowerThe rig draws watts around the clock if you want it available.Invisible until the electricity bill arrives.
CoolingEvery watt of compute becomes heat that has to leave the room.Folded into “the building” and never attributed.
Ops timePatching, monitoring, model updates, incident response: a fraction of a real person.The single biggest hidden cost, and the one that decides the outcome.
RedundancyWhat you spend to not be down when one box fails.The gap between “a nice demo” and “production”.
Setup (one-off)Initial integration, security, evaluation per workload.Real engineering effort before anything runs in production.

Side B: The API meter

Cost componentWhat it really is
UsageTokens in plus tokens out times the provider rate, summed over the month.
GrowthThe bill scales with adoption: every new use case adds to it, quietly.
FloorNear-zero fixed cost. You pay for what you use, including 0 EUR in a quiet month.

The payback formula

                   One-off cost to stand up self-hosting (hardware + setup)
Payback (months) = ───────────────────────────────────────────────────────────
                    Monthly API spend − Monthly self-host run-rate (opex)

Two honest consequences fall straight out of this formula:

  • If the monthly run-rate is close to the API spend, payback is never. Ops time alone can eat most of a small workload’s “savings”. For one light workload, self-hosting often does not pay back.
  • Payback collapses as you add workloads. Power, cooling, redundancy and (mostly) ops time are largely fixed per node. Put a second and third workload on the same hardware and the API side grows while the self-host side barely moves. Self-hosting is a consolidation play, not a single-app play.

That is the real thesis, and it is the honest one: self-hosting wins when you have enough AI work to keep one well-run node busy.

Part 2: Run the numbers yourself

All defaults below are the illustrative figures from the worked example and are clearly marked editable. Put in your own numbers and see where your tipping point lands. The faint dashed line shows the “hardware-only” math, deliberately there as the contrast to the honest line.

Illustrative figures, not an offer. Your numbers will differ.

Self-hosting (own hardware)

API (cloud LLM)

Self-host / month
API / month
Saving / month
Payback
Cumulative cost over 36 months
API (cloud) Self-host (full TCO) hardware-only: the easy (incomplete) math

The illustrative worked example

Setup (illustrative): a desk-scale node in the spirit of what we run, roughly 20,000 EUR of hardware (around 4 DGX-Spark-class boxes), suitable for a mid-size open-weight model.

Scenario 1: One workload (the cautionary case)

ItemIllustrative monthlyBasis
Hardware depreciation555 EUR20,000 EUR over 36 months
Power175 EURaround 1.2 kW average, 24/7, 0.20 EUR/kWh
Cooling55 EURaround 30 % overhead on IT power
Ops time1,100 EURaround 0.15 FTE of an MLOps/admin, loaded
Redundancy (amortised)150 EURspare-capacity / HA provision
Self-host run-ratearound 2,035 EUR / mo
Comparable API spendaround 3,000 EUR / morepresentative mid-size workload
Monthly differencearound 965 EUR / mo
One-off setup (integration, security, eval)around 10,000 EUR

Payback around 30,000 EUR / 965 EUR per month, so around 31 months. Not the six-to-seven months the naive “hardware vs. API bill” math promises. With full-TCO honesty, payback for a single light workload is slow, well over two years, and may not be worth it on cost alone. (Data residency and independence may still justify it.)

The widely quoted “around 6 to 7 month” payback (around 20,000 EUR hardware / around 3,000 EUR per month) is the hardware-only view. We show it side by side precisely to make the gap between the easy story and the honest one visible.

Scenario 2: Three workloads on the same node (the real case)

The ops, power, cooling and redundancy are largely shared; only modest setup is added per workload.

ItemIllustrative monthly
Self-host run-rate (shared node, 3 workloads)around 2,400 EUR / mo
Comparable API spend (3x around 3,000 EUR)around 9,000 EUR / mo
Monthly differencearound 6,600 EUR / mo
One-off setup (hardware + 3x integration)around 40,000 EUR

Payback around 40,000 EUR / 6,600 EUR per month, so around 6 months. Same hardware, same operating discipline, but now the consolidation does the work. This is where self-hosting earns its place.

Part 3: The honest footnote

We run this ourselves, so we will tell you the part the calculator cannot: operating it is the project. Opteria runs exactly one node with one GPU plus a staging environment: small, but real. We do not sell borrowed scale; we sell the method and the outcome. The reason our number for “ops time” is not zero is that we pay it every month, and we would rather you budget for it honestly than discover it after the hardware arrives.

So use the calculator for the cost story. Then ask the harder questions it does not price: Do you have enough AI workload to fill a node? Who owns uptime at 2 a.m.? How sensitive is the data? Those decide self-host vs. cloud at least as much as the euros do, and they are exactly the per-workload decision we make at Opteria.

Related: our broader Cloud-LLM vs. own infrastructure cost calculation and the decision tree: self-host or cloud.

Ready to implement AI in production?

We analyse your process and show you in 30 minutes which workflow delivers the highest ROI.