When should you stop renting tokens? The build-vs-buy crossover for internal RAG

4 minute read

Published:

中文版 Chinese version

Every company onboards Anthropic or OpenAI the same way: an engineer wires up the API in an afternoon, the demo sings, and six months later finance forwards a token bill nobody budgeted for. It’s the AWS story replayed — rent for elasticity, then discover at scale that you’re renting what you should own.

The question is where “scale” actually starts. I priced it.

The three tiers, with current numbers

All prices verified September 2026. Blended rates assume a 4:1 input-to-output ratio, typical for input-heavy RAG and agent workloads — and note the trap the field guide crowd keeps flagging: output tokens cost 5–8x input, so comparing input prices alone picks the wrong model.

TierReference priceBlended $/M tokens
Frontier API (Claude Sonnet 5: $3 in / $15 out)Anthropic pricing$5.40
Frontier API (OpenAI GPT-5: $1.25 in / $10 out)pricepertoken.com$3.00
Managed open 70B (Together AI Llama 3.3 70B: $1.04 / $1.04)together.ai/pricing$1.04
Self-host 32B-class on 1×H100 ($2.70/hr on-demand)RunPod/Lambda pricing~$0.50 marginal

Side note: OpenAI cut GPT-5.6 Sol to $4/$20 for three months starting August 21 (Reuters via quasa.io). The price war is real, and it only moves the crossover points — it doesn’t remove them.

Build-vs-buy crossover: monthly cost vs token volume, and annual cost at 10B tokens/month

The crossover points

Assumptions, stated plainly: one H100 at $2.70/hr costs $1,972/month and sustains ~1,500 tokens/sec on a 32B-class model (conservative — short-context serving benchmarks run 2–4x higher), i.e. ~3.9B tokens/month per GPU. A platform team of 2–3 engineers runs ~$1.5M/year fully loaded at Staff-level comp.

  • Raw infra vs frontier API: ~365M tokens/month. That’s roughly $2K/month in API spend. At 100M tokens per employee per month — normal once agents are in the loop — that’s about 4 employees. At 10M (heavy chat), ~37 employees. The raw crossover is shockingly small, and it’s the number vendors hope you never compute.
  • Managed open APIs vs self-host: ~1.9B tokens/month. The managed tier ($1.04/M, zero ops) captures most of the savings. Self-hosting only beats it once you’re filling GPUs.
  • With a dedicated platform team: ~23.5B tokens/month vs API, ~122B vs managed open. This is the crossover that matters, and it’s two orders of magnitude higher than the raw-infra one. The GPU is cheap; the engineers aren’t.

Panel B makes it visceral at 10B tokens/month: $648K/yr (Sonnet 5) vs $360K (GPT-5) vs $125K (managed open) vs $71K (self-host marginal) vs $1.57M (self-host + team). The team tax dominates everything.

What to do Monday morning

  1. Compute your blended $/M, not your input $/M. Pull last month’s in/out split. If you’re comparing $3 vs $1.25 input prices while your workload is output-heavy, you’re doing the arithmetic wrong.
  2. If you’re past ~350M tokens/month, price the managed open tier this week. Together/Fireworks-class 70B serving at ~$1/M is 3–5x cheaper than frontier APIs with zero operational burden. For most companies, this is where the journey should stop.
  3. Self-host only with sustained volume and an existing platform team. Past ~2B tokens/month with engineers already on payroll, self-hosting wins. Hiring a dedicated team for it needs ~20B+/month sustained to pencil out.
  4. Build the showback dashboard regardless. When internal tokens cost $0, nobody effort-routes, caches, or distills — the cost converts into capacity contention everyone pays. A monthly “your team burned $X at API rates” number, with no actual charge, recovers the incentive to be efficient.

Caveats, stated plainly

  • Quality parity is assumed and it’s false at the top end. A 70B open model is not Sonnet 5 on hard reasoning — that’s exactly what the frontier premium buys. The answer is routing: frontier for the hard tasks, cheap tiers for the bulk. The gateway from Monday’s post, again.
  • Self-host lines assume filled GPUs. At 50% utilization your effective $/M doubles; internal tools spike 9-to-5 and idle at night. The API premium is elasticity insurance — the same reason most companies never leave AWS.
  • Throughput is the sensitive variable. If your workload sustains 3,000 tok/s per GPU instead of 1,500, every self-host breakeven halves. Recompute with your own serving numbers before committing.
  • Compliance and data residency can force self-hosting regardless of cost. Regulated data doesn’t care about your crossover chart.

Thesis in one line: stop asking “API or GPUs” — the real build-vs-buy question is the team, not the hardware, and for most companies the answer is the managed open tier in the middle.

Sources