← Back

Buy tokens, rent GPUs, or own the rack?

Until recently, self-hosting a frontier model wasn’t a serious question. Open-weight models sat a generation behind the proprietary ones, so there was no calculation to make.

Then the gap closed. Not to zero, but to about thirty Elo. Kimi K3 shipped with open weights and sits second on the frontend code arena at 1,682, behind Claude Opus 5 Max on 1,712 and ahead of everything else1. Its weights fit, on paper, on a single eight-GPU node. You can download the model tonight for nothing.

The box is another matter. Assuming you can make the capital investment. And assuming, once you’ve done the arithmetic, it turns out to be worth doing at all.

So I modelled four ways to run the same workload: a rack in your office, owned hardware in a colocation facility, rented dedicated GPUs, and pay-per-token APIs. I’ll call that last one the meter throughout.

The answer isn’t “buy the rack”. Under about 44 engineers the meter stays cheapest. Between roughly 44 and 95, rented GPUs win. Above that owning starts to pay, and by 750 engineers it beats the meter roughly six to one.

The number that decides it is utilisation. A rack bought on a three-year schedule costs the same at three in the morning whether it’s serving tokens or warming the building.

K3 is only the worked example here. The same method applies to whatever ships next.

Four ways to buy inference

OptionYou ownYou rentYou pay for
Office rackGPUs, power infrastructure, coolingnothingCapex, electricity, an electrician, floor space, people
Colo rack (your hardware, someone else’s building)GPUsSpace, power, coolingCapex, rack rental per kW2, people
Dedicated / rented GPUsnothingWhole GPUs by the hourReserved hours, people
The meter (per-token API)nothingnothingTokens

As you move down that table your fixed commitment falls and your flexibility rises. What you give up is control, and at sufficient scale, unit economics. Finding where that trade flips is the whole exercise.

Before the spreadsheet, though, a more basic question: can you actually put one of these machines in an office?

Can we?

In season two of Silicon Valley, Gilfoyle’s servers catch fire. The team piles too much load onto Anton, it maxes out the amperage, blows the main breaker and burns. Everyone remembers it as a gag about hubris. Here, we’re treating it as commentary on electrical engineering.

Pied Piper's servers burning in the background of the Hacker Hostel

Anton, moments after the main breaker gave up. Silicon Valley, HBO.

Apologies if you haven’t seen Silicon Valley. It’s a modern classic, ahead of its time.

K3 is 2.8 trillion parameters, 104 billion active per token, 896 experts, quantised to MXFP4 straight out of training3. That is 1,390GB of weights on paper. An eight-GPU B200 node gives you 1,440GB, so on paper it fits with 50GB to spare. Reports of the actual checkpoint put it nearer 1.56TB once everything that is not a quantised weight is counted, in which case it does not fit at all. This is the r/selfhosted wet dream: the second-best coding model on earth, humming away in a cupboard, yours.

Then you read the spec sheet.

Spec, one DGX B200ValueWhat it means in an office
Power draw14.3 kW4Two full induction hobs, every ring, permanently
UK ring main7.4 kWOne server wants two entire circuits to itself
Heat output~48,800 BTU/hrThree to four domestic aircon units, running flat out
Airflow2,145 CFM4Not a cupboard. A plant room
Weight130 kg4Before the rack, PDUs and cooling
Two nodes per rack2x 380V three-phase 32A4You are calling an electrician, not buying an extension lead

One inconsistency to own: the power, weight and airflow figures are NVIDIA’s for the DGX B200, its own integrated appliance, while the $450,000 I model is a street price for an 8-GPU HGX B200, the board an OEM builds a server around. A DGX lists nearer $515,000. I use the cheaper number with the pricier machine’s power draw, which flatters self-hosting on capex and penalises it on electricity.

Anton didn’t blow that breaker because a writer wanted a fire in act three. It blew because that’s what happens when you hang industrial load off domestic wiring.

One node is the charitable case, by the way, and it’s charitable in a second way too. Those 50GB of headroom leave nothing for KV cache, the working memory a model needs to hold a long conversation, so a real production deployment needs more accelerators than this. Moonshot recommends 64 or more5. That’s eight of these boxes, 114 kilowatts, in your office. Less a server room than a micro nuclear reactor waiting for someone to announce a Series A around it.

I’ve modelled the single node anyway, because it is the cheapest thing that could plausibly hold the weights. Be careful with that, though: it is not a worst case. Staff and electrical work are close to fixed, so more nodes spread them and the economics improve. A one-node result is the hardest test self-hosting faces, not the easiest.

Then there’s supply. Blackwell allocation goes to the biggest buyers first, so getting hold of one means queueing behind Musk.

You need a Gilfoyle

Gilfoyle is the systems guy. He owns the racks, trusts nothing, and he’s the only one in the building who knows why the cluster is on fire.

Which brings us to my favourite fiction in a self-hosting business case: a fraction of an engineer.

You can’t hire 0.3 of a person. The role here runs vLLM or SGLang in production, knows what expert parallelism is, debugs NCCL at two in the morning, and keeps several hundred gigabytes of mixture-of-experts hot and serving. Robert Half puts the median London machine learning engineer at £102,000 and the upper quartile near £119,0006. I’ve modelled £120,000, deliberately an upper-quartile hire, because the median one can’t do this job.

They will also have opinions. About your cloud spend, about the keyboard they require, about the equity owed to the one person in the building who understands the thing everyone now depends on. That is the real cost of running it yourself, and it appears on no GPU spec sheet.

You need two of them, because one is a single point of failure with a passport and strong feelings about annual leave.

LineValue (GBP)
Base salary (upper quartile)£120,000
Overhead multiplier (employer NI, pension, kit)1.3
Loaded cost per person£156,000
People required2
Annual total£312,000 (about $396,000)

That number doesn’t depreciate. It doesn’t get cheaper, and it doesn’t care whether your hardware is busy.

The workload decides more than the hardware

Most self-hosting sums I’ve seen get costed against chat, which is the worst possible case for owning hardware and nothing like what a development team does all day.

Agentic coding runs at 4.17 million tokens per task, with an input-to-output ratio of roughly 153:17. Nothing here matters more than that ratio except utilisation.

On your own hardware it helps. Prefill, the reading of your prompt, and decode, the writing of the answer, are different jobs with different costs. vLLM clocks 26,200 tokens per GPU per second on prefill against 10,100 on decode8, so prefill is two and a half times more efficient. Agentic coding is almost entirely prefill. The work your team does happens to be the work GPUs are best at.

On the meter it helps too, because input tokens are five times cheaper than output.

ComponentWorking$/1M tokens
Uncached input0.30 × 0.9935 × $51.49
Cached input (reads at 0.1×)0.70 × 0.9935 × $5 × 0.10.35
Output0.0065 × $250.16
Blended, 70% cache hits$2.00
Blended, no caching at all0.9935 × $5 + 0.0065 × $25$5.13

That 70% is an assumption, not a measurement, and it is the second most important number in this post. It more than halves the meter, from $5.13 to $2.00. Halve the hit rate to 35% and the meter runs $3.57, which moves every crossover below sharply toward self-hosting. Go the other way and the meter wins nearly everywhere.

Two more things about that $2.00 before you trust it. Anthropic charges roughly 1.25x input to write a cache entry, and I have priced every uncached token as a plain read. If all of the uncached 30% were writes, the meter is $2.37 rather than $2.00. The truth is somewhere between.

Someone will ask why I price self-hosted K3 against Opus 5 rather than against K3’s own API, which is cheaper at $3 and $15 per million with cache hits at $0.309, about $1.20 on this token mix. Because for most of the companies this post is aimed at, sending production traffic to a Chinese cloud is not a live option, whatever the price list says. That is not a technical judgement, it is a procurement one, and it gets made above your head. The realistic choice is running K3 yourself or buying Opus 5 from Anthropic, which is what the tables compare.

It is also, incidentally, the same instinct that makes people self-host in the first place. If you were relaxed about where the weights run, you would use the API and stop reading.

So before you use any of this: go and look at your actual cache hit rate. It matters more than the price of a GPU.

What one box actually serves

To compare four options I need two numbers: how much work one node processes, and how much work one engineer generates. Neither is exact, so here is every assumption, including the ones that ruin the fantasy.

StepWorkingResult
Input share of tokens153 ÷ 1540.9935
Blended throughput per GPUharmonic mean of 26,200 prefill and 10,100 decode at that mix25,932 tok/s
Derate for K3 vs DeepSeek-class104B active params vs ~37B× 0.35
Derate for B200 vs GB200benchmark ran on GB200× 0.70
Derate for real traffic vs benchmarkbenchmarks are tidy, production isn’t× 0.60
Derated throughput per GPU25,932 × 0.1473,812 tok/s
Per node× 8 GPUs30,496 tok/s
Annual capacity at 100% duty× 31,536,000 seconds962,000M tokens

Against that, demand:

StepWorkingResult
Tasks per engineer per dayworkload assumption, not measured6
Tokens per taskBai et al.74.17M
Working days per year260 less leave and bank holidays230
Tokens per engineer per year6 × 4.17M × 2305,755M

Flat out, one node serves about 167 engineers. At a 21% annual duty cycle, roughly office-hours usage, it supports about 35.

The exact capacity figure is arguable and I’ll come back to how wrong it might be. Once the hardware is bought, utilisation swamps every other variable.

Utilisation decides it

Owned hardware costs the same asleep as it does flat out. The depreciation runs, the rack rent runs, and your two Gilfoyles get paid regardless. Electricity is the only line that tracks demand, and electricity turns out to be pocket change.

UtilisationOffice rack $/1MColo $/1MMeter $/1M
5%12.2312.812.00
10%6.146.402.00
21% (office hours only)2.953.052.00
35%1.791.832.00
50%1.271.282.00
80%0.810.802.00
100%0.660.642.00

The office line crosses the meter at 31.2% utilisation, colo at 32.0%.

One caveat that belongs right here rather than at the end: that is a one-node curve. Two platform engineers and the electrician’s invoice do not scale with node count, so on two nodes the crossover falls to about 21% and on five to about 14%. The bigger you get, the less utilisation you need to justify owning.

Now hold a real team up against it. Eight hours a day, 230 days a year, works out at 1,840 hours from a possible 8,760. A 21% duty cycle. A team that works office hours and then goes home sits below break-even, and self-hosting loses.

So the question was never whether open-weight models are good enough, or whether GPUs are cheap enough. Both are settled. The question is whether you can fill the night shift with batch jobs, evals, reindexing and the CI nobody watches. Get the box to 35% and you’re winning. At 50% you’re winning comfortably. Let it idle every evening and you’ve bought a very expensive space heater on a three-year lease.

Independent analyses of Azure’s provisioned throughput put the break-even against pay-as-you-go around 80% sustained utilisation10. Neither AWS nor Microsoft publishes a break-even figure of its own, so treat that as third-party arithmetic rather than vendor guidance. That is nowhere near my 31%, and I am not going to pretend otherwise. They are measuring different things: theirs is a margin-bearing product priced against its own list, mine is raw cost against a rival’s list. What the two share is direction, which is that reserved capacity needs to be busy most of the time before it pays.

Rented GPUs win the middle

Utilisation sets the unit price. Team size decides which option wins, because that is what spreads the fixed costs.

At 25 engineers, one node sits at 15% duty and the meter wins outright. Two platform engineers cost $396,240 a year. The entire token bill they would be replacing is $287,777. The staff cost more than the problem. Nothing else on the spreadsheet is worth reading.

At 60 engineers, one node runs at 36% duty and rented capacity edges ahead.

Annual line (USD)OfficeColoRentedMeter
Hardware depreciation11150,000150,00000
Power infrastructure25,400000
Electricity31,821000
Rack rental069,56400
GPU rental00138,3820
Platform engineers396,240396,240396,2400
Token charges000690,664
Total603,461615,804534,622690,664
Per engineer10,05810,2648,91011,511
Per 1M tokens1.751.781.552.00

Note what is missing from that table: a staffing discount for renting. Every self-hosted option carries two whole platform engineers, because you cannot page a fraction of a person at two in the morning whoever owns the metal. Renting saves you the capex, the electrician and the queue behind Musk. It does not save you the payroll.

Which is why the win is narrow. Cheapest to dearest across all four options here is 29%, comfortably inside the error bars on my own assumptions.

I went looking for someone who had costed this option against the other three and came up short, which puzzles me, because there is nothing obscure about renting GPUs. Lambda, CoreWeave, Spheron and a dozen others publish B200 hourly rates, and one aggregator tracks more than 26 of them12. The middle is very much for sale, and for a 60-engineer team it beats both ends.

By 750 engineers, five nodes run at 90% duty and renting loses badly, because it bills by the hour and at that utilisation you are paying nearly all of them. Owning the hardware costs $1,494,059 in a colo or $1,565,997 in your own office, against $2,126,018 rented and $8,633,301 on the meter. This is the one place in the whole model where owning is obviously right, and it beats the meter roughly six to one.

Notice how close the office and colo columns are, though. Above about 95 engineers they run within a few percent of each other and swap places depending on node count. Given that my office option pays no rent, the honest reading is that they are the same answer, and the choice between them is about who you want changing the filters.

Team sizeCheapest optionAnnual cost per engineerMain reason
25The meter$11,511Platform staff cost more than the entire token bill
60Rented GPUs$8,910Enough duty to want capacity, not enough to justify capex
750Owned hardware (colo)$1,992Near-continuous demand makes rental hours expensive

Tokens aren’t money

There’s a persistent assumption in technology planning, and I say this as someone unmistakably part of the problem, that tokens are the new barrel of oil. The universal unit of exchange. Whole companies now run planning cycles denominated in them.

Fully loaded and fully utilised, the owned node in this model costs about $0.66 per million tokens against $2.00 on the cached meter. Electricity alone is 6.6 cents of that, which is a tenth of the total and, annoyingly, a factor of ten off the number above it. Both are right.

LineWorkingValue
Node power14.3 kW × 8,760 hours125,268 kWh/yr
At UK business rates13× £0.25£31,317
With office cooling overhead (PUE 1.6)× 1.6£50,107
In dollars× 1.27$63,636
Tokens produced at full duty962,000M
Electricity per 1M tokens6.6 cents

The point isn’t that anyone should charge seven cents. A token’s price includes hardware, staff, risk, research and margin, and it should. The point is that a token is a retail billing abstraction over GPU-seconds, with an exchange rate set by whoever is selling it. If tokens really are the new oil then the hyperscalers are OPEC, with the useful advantage that they get to revise the physics whenever margins need help.

Which is where this goes next, and it’s less a prediction than a pattern that has already run once. AI compute gets financialised exactly like Web 2.0 compute did. EC2 launched priced by the instance-hour, and within a few years the real market was reserved instances, savings plans, spot and sustained-use discounts. Large buyers stopped paying the on-demand rate years ago.

It’s already starting. Bedrock sells provisioned throughput by the model unit-hour, from $4.11 to $49.50 depending on model, and the rate drops the longer you commit10. Azure sells the same shape as provisioned throughput units. Those aren’t token prices, they’re capacity reservations with a commitment curve.

The tell is what’s missing from that page. No current Claude model has a published provisioned-throughput rate at all. The reserved market for frontier models exists, but it is quoted rather than listed, which tells you how early this is.

Follow it forward and you stop buying tokens. You reserve capacity for your baseline, take the sustained-use discount, and burst to the meter when you spike. Tokens turn into exhaust from a pipe you have already paid for, and the thing you budget in becomes the GPU-hour, which at least has a real supply curve behind it.

Should we?

Some of this might read like an argument for going out and buying some metal. It isn’t, quite.

You would be buying a depreciating asset while the meter price falls. Opus 5 launched at exactly Opus 4.8’s price, $5 and $25 per million tokens, and it is a materially better model14. The price list didn’t move while the thing behind it improved. So your capex is a three-year bet that frontier intelligence gets cheaper more slowly than your hardware wears out, and the frontier has been moving on a timescale of months. Hardware you spec today is stuck with today’s economics until 2029.

Both things are true at once, then. Tokens carry a fat margin, and buying machines is a poor way for most people to escape it. The way out is that you don’t have to buy a rack to stop paying retail. Reserve the capacity, don’t buy the building.

Two things survive the arithmetic, though only one of them cleanly.

Sovereignty. If your data can’t leave the building then you aren’t comparing against $2.00, you are comparing against nothing, because the meter isn’t for sale to you at any price. Run the utilisation numbers anyway, but you have already decided.

Latency, but less than you’d hope. A model in the same rack saves you a network round trip. That’s perhaps forty milliseconds set against the thirty-odd seconds the thing spends thinking, so for agentic work it’s noise. It earns its keep on the tight stuff, autocomplete and inline suggestions, where you’re chasing a sub-100ms budget and the network is a real share of it. If that’s not your workload, don’t put it in the business case.

So here’s what I’d actually do: stop picking one of the four. Reserve enough capacity to cover your baseline and meter the rest. Predictable high-volume work runs overnight on the pipe you’ve already paid for, at the utilisation where it wins two to one or better. The hard reasoning goes to the frontier model on the meter, where you’re paying for something you can’t reproduce and can’t depreciate.

In practice, placement rather than model choice is the main economic lever. That was the argument I made about routing between models, one layer further down the stack.

Where this model could be wrong

Four things.

My throughput derate is probably too harsh, and correcting it favours renting. I knocked the vLLM benchmark down by a factor of seven to cover K3 being bigger, the hardware being air-cooled and real traffic being messier than a benchmark. That’s napkin maths, and a rough check against the raw FLOPs says the true figure could be three or four times higher. Rerun at that number and 200 engineers flips from owning to renting, because faster hardware cuts rented GPU-hours immediately while doing nothing about the capex you already committed. It moves against a rack, not for one.

Tasks per engineer per day is a guess. I used six. It drives absolute demand and therefore the team-size results. It doesn’t touch the utilisation table, which is why that is the stronger claim and why I put it first.

The single-node configuration is optimistic, as flagged earlier. Moonshot wants 64 accelerators for production and I modelled eight, which is the most generous reading self-hosting could ask for.

The office option gets its building for free. It pays for hardware, the electrical and cooling install, and the electricity. It does not pay rent, business rates, insurance, fire suppression, UPS, PDUs, structural work for 130kg a node, or a DNO supply upgrade, and above about two nodes a UK commercial building will need that upgrade rather than an electrician. Colo pays for all of it inside the £285/kW/month. That gap is part of why office beats colo everywhere in my tables, and you should treat the two as closer than they look. Charging the office option half of colo’s all-in rate still leaves it ahead at every team size, which is the only reason I have not tried to invent a number for it.

Every number has a source and a confidence rating in the workbook behind this. The low-confidence ones are the three you have just read.

So the arithmetic might be wrong at the edges. The physics isn’t. Fourteen kilowatts is fourteen kilowatts whether my throughput estimate holds or not, the heat has to go somewhere, and the electrician still wants paying.

I could buy tokens and never think about any of this again. Or I could think about all of it, and get heatstroke wiring up another graphics card.

Footnotes

  1. Arena WebDev leaderboard, 28 July 2026, 492,170 votes. claude-opus-5-max 1,712; kimi-k3-max 1,682; claude-opus-5-high 1,669; claude-fable-5 1,628; gpt-5.6-sol-xhigh 1,623; claude-opus-4-8-thinking 1,568. Worth holding onto: K3’s coding rank flatters it. On the main text leaderboard (27 July 2026) kimi-k3-max is eleventh on 1,486 ±10, behind claude-fable-5 (1,508) and claude-opus-5-max (1,495), and in a statistical tie with claude-opus-4-8-thinking (1,484 ±5). It is a coding specialist, not the best model in the world.

  2. I model colo at £285 per kW per month. I pulled that number out of my arse, because I could not find a source good enough to stand behind. For what it is worth: UK wholesale colocation runs £150-300 per kW per month at 1MW and above, retail London is quoted at £300-800 per cabinet rising past £2,000 for density, and high-density GPU racks carry a premium nobody will put a number on. My deployment is 14 to 72kW, far too small for wholesale pricing. If you know what a 30kW GPU rack actually costs in London, tell me and I will correct this. I treat the rate as all-in, covering space, cooling, resilience and power, because charging metered electricity on top of it would double-count the largest line on a colo bill.

  3. Moonshot AI, Kimi-K3 model card. Hugging Face. 2.8T total parameters, 104B active (16 of 896 experts), MXFP4 weights with MXFP8 activations, 1,048,576 token context.

  4. NVIDIA, Data Center Best Practices with DGX B200. Spec sheet. Estimated system power 14.3kW max, weight >130kg, and 150 CFM per kilowatt, giving 2,145 CFM per node (the document quotes 4,290 CFM for a two-node rack). Two nodes per rack require 2x 380V three-phase 32A feeds. The heat figure is my own conversion, not NVIDIA’s: 14.3kW x 3,412 = ~48,800 BTU/hr. 2 3 4

  5. Moonshot AI, Kimi K3 launch post: “we recommend deploying Kimi K3 on supernode configurations with 64 or more accelerators”. A recommendation, not a floor. Moonshot publishes no minimum GPU count.

  6. Robert Half, 2026 Salary Guide, machine learning engineer, London. 50th percentile £102,000, 75th percentile approximately £119,000.

  7. Bai et al. (2026), How Do AI Agents Spend Your Money? Paper, Figure 1. Agentic coding averages 4.17M tokens per task at an input-to-output ratio of 153.85:1, against 1.33 for code chat and 0.16 for code reasoning. The paper also reports $1.86 per task, but that is an average across eight models, most far cheaper than Opus 5, and works out at roughly $0.45 per 1M tokens. I take the token volume and the ratio from the paper and price them myself: the same task on Opus 5 at a 70% cache hit rate is about $8.34. Do not read the two cost figures as comparable. 2

  8. vLLM, Driving vLLM WideEP and Large-Scale Serving Toward Maturity on Blackwell (Part I), 3 February 2026. Link. 26.2K prefill and 10.1K decode tokens per GPU-second on GB200 for DeepSeek-style MoE at 2K in / 2K out. Two things to carry with the number: it comes from a disaggregated 16-GPU topology (four prefill instances of two GB200 each, one decode instance of eight), and throughput plateaus around a 64K batch size. Scaling it linearly to a single air-cooled eight-GPU box is exactly the mistake my derates exist to cover.

  9. Kimi K3 on Moonshot’s API is $3 per million input tokens and $15 per million output, with cache hits at $0.30, flat across the full 1M context. Rates are widely reported but I have not been able to confirm them on Moonshot’s own pricing page, so treat them as indicative. On this post’s token mix that blends to $1.20 per million against $2.00 for Opus 5.

  10. AWS, Bedrock pricing. Published provisioned-throughput rates span $4.11 to $49.50 per model unit-hour and cover legacy models only (Llama 2 $21.18 for one month, Cohere Command $49.50 for one month against $23.77 for six, Command Light from $8.56 down to $4.11). No current Claude model carries a published provisioned rate. Azure’s equivalent is the provisioned throughput unit, whose mechanics are documented at Microsoft Learn though the page carries no dollar figures. The ~80% break-even is third-party analysis, not vendor guidance, and neither platform publishes such a number. 2

  11. Street pricing for an 8-GPU HGX B200 clusters around $400,000 to $500,000 and I have modelled $450,000 over a three-year straight-line life with no residual value. Treat that as my assumption rather than a citation: the only public sources are vendor content-marketing pages with no named author and no primary references, one of which concedes that no functioning secondary market exists to verify against. If you have a real quote, yours beats mine.

  12. Published B200 hourly rates vary enormously by provider and commitment, roughly $2.14 to $27.04 per GPU-hour. Lambda lists $4.99 and Spheron $3.70 on demand against $2.74 spot, while AWS p6 runs $14.24. getdeploying aggregates more than 26 providers. I have modelled $5.50, which sits between neocloud and hyperscaler pricing and matches neither, so treat it as a midpoint assumption rather than a quote.

  13. DESNZ, Quarterly Energy Prices, June 2026, tables 3.4.1 to 3.4.2. Volume-weighted average non-domestic electricity price of 24.14p/kWh in Q1 2026, provisional, including Climate Change Levy and excluding VAT. I have rounded up to 25p, which makes the self-hosted options look slightly worse rather than better. PUE is power usage effectiveness, the ratio of total facility power to the power actually reaching the equipment.

  14. Anthropic, Models overview. Opus 5 and Opus 4.8 both list at $5 per million input tokens and $25 per million output tokens.