When should you stop renting tokens? The build-vs-buy crossover for internal RAG
Published:
Every company onboards Anthropic or OpenAI the same way: an engineer wires up the API in an afternoon, the demo sings, and six months later finance forwards a token bill nobody budgeted for. It’s the AWS story replayed — rent for elasticity, then discover at scale that you’re renting what you should own.
The question is where “scale” actually starts. I priced it.
The three tiers, with current numbers
All prices verified September 2026. Blended rates assume a 4:1 input-to-output ratio, typical for input-heavy RAG and agent workloads — and note the trap the field guide crowd keeps flagging: output tokens cost 5–8x input, so comparing input prices alone picks the wrong model.
| Tier | Reference price | Blended $/M tokens |
|---|---|---|
| Frontier API (Claude Sonnet 5: $3 in / $15 out) | Anthropic pricing | $5.40 |
| Frontier API (OpenAI GPT-5: $1.25 in / $10 out) | pricepertoken.com | $3.00 |
| Managed open 70B (Together AI Llama 3.3 70B: $1.04 / $1.04) | together.ai/pricing | $1.04 |
| Self-host 32B-class on 1×H100 ($2.70/hr on-demand) | RunPod/Lambda pricing | ~$0.50 marginal |
Side note: OpenAI cut GPT-5.6 Sol to $4/$20 for three months starting August 21 (Reuters via quasa.io). The price war is real, and it only moves the crossover points — it doesn’t remove them.

The crossover points
Assumptions, stated plainly: one H100 at $2.70/hr costs $1,972/month and sustains ~1,500 tokens/sec on a 32B-class model (conservative — short-context serving benchmarks run 2–4x higher), i.e. ~3.9B tokens/month per GPU. A platform team of 2–3 engineers runs ~$1.5M/year fully loaded at Staff-level comp.
- Raw infra vs frontier API: ~365M tokens/month. That’s roughly $2K/month in API spend. At 100M tokens per employee per month — normal once agents are in the loop — that’s about 4 employees. At 10M (heavy chat), ~37 employees. The raw crossover is shockingly small, and it’s the number vendors hope you never compute.
- Managed open APIs vs self-host: ~1.9B tokens/month. The managed tier ($1.04/M, zero ops) captures most of the savings. Self-hosting only beats it once you’re filling GPUs.
- With a dedicated platform team: ~23.5B tokens/month vs API, ~122B vs managed open. This is the crossover that matters, and it’s two orders of magnitude higher than the raw-infra one. The GPU is cheap; the engineers aren’t.
Panel B makes it visceral at 10B tokens/month: $648K/yr (Sonnet 5) vs $360K (GPT-5) vs $125K (managed open) vs $71K (self-host marginal) vs $1.57M (self-host + team). The team tax dominates everything.
What to do Monday morning
- Compute your blended $/M, not your input $/M. Pull last month’s in/out split. If you’re comparing $3 vs $1.25 input prices while your workload is output-heavy, you’re doing the arithmetic wrong.
- If you’re past ~350M tokens/month, price the managed open tier this week. Together/Fireworks-class 70B serving at ~$1/M is 3–5x cheaper than frontier APIs with zero operational burden. For most companies, this is where the journey should stop.
- Self-host only with sustained volume and an existing platform team. Past ~2B tokens/month with engineers already on payroll, self-hosting wins. Hiring a dedicated team for it needs ~20B+/month sustained to pencil out.
- Build the showback dashboard regardless. When internal tokens cost $0, nobody effort-routes, caches, or distills — the cost converts into capacity contention everyone pays. A monthly “your team burned $X at API rates” number, with no actual charge, recovers the incentive to be efficient.
Caveats, stated plainly
- Quality parity is assumed and it’s false at the top end. A 70B open model is not Sonnet 5 on hard reasoning — that’s exactly what the frontier premium buys. The answer is routing: frontier for the hard tasks, cheap tiers for the bulk. The gateway from Monday’s post, again.
- Self-host lines assume filled GPUs. At 50% utilization your effective $/M doubles; internal tools spike 9-to-5 and idle at night. The API premium is elasticity insurance — the same reason most companies never leave AWS.
- Throughput is the sensitive variable. If your workload sustains 3,000 tok/s per GPU instead of 1,500, every self-host breakeven halves. Recompute with your own serving numbers before committing.
- Compliance and data residency can force self-hosting regardless of cost. Regulated data doesn’t care about your crossover chart.
Thesis in one line: stop asking “API or GPUs” — the real build-vs-buy question is the team, not the hardware, and for most companies the answer is the managed open tier in the middle.
Sources
- Anthropic Sonnet 5 pricing ($3/$15 per 1M, standard from Sep 1 2026): worthview.com
- OpenAI GPT-5 pricing ($1.25/$10 per 1M): pricepertoken.com
- GPT-5.6 Sol promotional cut ($4/$20 through Nov 21 2026): quasa.io
- Together AI serverless pricing (Llama 3.3 70B $1.04/$1.04 per 1M): together.ai/pricing
- H100 on-demand rental ($2.59–$3.29/hr RunPod/Lambda, Sep 2026): getdeploying.com
- Output-token cost trap and caching multipliers: dev.to field guide
Share on
Twitter Facebook LinkedIn☕ Buy me a coffee! 💝
If you found this article helpful, consider buying me a coffee to support my work! 🚀
