The Call Center Doctors rented four Nvidia H200s to swap DeepSeek in for Claude. Their coding agents re read 96% of every prompt, and DeepSeek's caching tier was already priced for that.
A 4-GPU server rented to escape a cloud AI bill ended up costing The Call Center Doctors more than twice its existing Anthropic subscription. It also could not undercut the cheapest API for the same work. The firm's write-up puts the headline number, a $13,200 monthly projection, on top of a roughly three-hour test that actually cost $9.19 an hour in spot capacity, with the non-spot rate of $18.37 an hour doing the extrapolating.
The extrapolation matters because the workload did not behave the way a self-hosting pitch assumes. In one-minute isolated tests, the box moved 16,621 new input tokens a second, 521,027 cached input tokens a second, and 5,281 output tokens a second, per the firm's Hugging Face dataset. In a separate mixed run with 48 agents, it moved 11,196 new input, 131,904 cached input, and 521 output tokens a second.
The asymmetry is the story. Coding agents like Claude Code resend the full conversation on every step, so most of what the model sees is text it has already read. The Call Center Doctors' own production mix, 41.6 new input tokens and 1,042 cached input tokens per output token, put the read/write ratio at roughly 25-to-1 in favor of cached re-reads. DeepSeek's API prices cached input at a steep discount precisely because it expects that shape. The rented box spent most of its capacity re-reading and still billed the firm for idle hours at the on-demand rate.
That is how the monthly number crossed $13,200. At the stated non-spot rate, four H200s come to about $13,200 a month of continuous operation, even though the test only ran for three hours. By comparison, the firm reports about $5,500 in Claude subscriptions and 5,610 merged changes over September 1–27 on one server. Its own counterfactual: identical token usage through DeepSeek's API would have run roughly $3,500 off-peak or $7,000 at peak. The box could not beat either number on its own terms.
A workload that mostly re-reads is the API's home turf. A workload that mostly writes new tokens is where the rented box can pay for itself. Long-context generation, bulk code synthesis from a small prompt, and batched document drafting tilt toward the box. Conversational agents, multi-step tool use, and any loop that re-sends prior context tilt toward the API. Cache-hit ratio, not GPU list price, is the right metric to measure before renting.
The firm flagged a separate concern: sandbox escape worries kept DeepSeek's builder agents offline for the test. Read-only reviewer agents scanned 2,377 folders and filed 32 bug reports. Zero DeepSeek-authored changes were merged. The security finding is real and stands on its own. It does not establish that DeepSeek cannot code. It establishes that the operator did not let it.
For anyone weighing the same trade-off, the lesson is concrete: measure the cache-hit ratio first. The Call Center Doctors' full write-up, including the test scripts and raw numbers, is published on its site.