1. Model Strategy & Token Traffic
Variables
1,200,000
2. Fine-Tuning & GPU Infrastructure
OpEx / CapEx
4 GPUs
40 hrs ($4.8k)
ℹ️ vLLM Throughput Assumption: Optimized with PagedAttention continuous batching delivering ~1,400 completed tokens/sec across the GPU cluster.
Monthly Cloud API Burn
$9,000
1.44B input · 540M output
Monthly GPU Hosting OpEx
$7,270
4x H100 @ $2.49/hr · 730h
One-Time Setup CapEx
$4,880
Compute: $80 · Labor: $4,800
Crossover Break-Even Vol
969,000
Requests / month needed
Cumulative 12-Month Financial Spend
API Token Spend
Self-Hosting Total
Month 0 (Setup)
Month 3
Month 6
Month 9
Month 12
Unit Economics per 1,000 Invocations
Normalized
API Cost / 1k Req
$7.50
Self-Hosted / 1k Req
$6.06
Unit Delta
-19.2%