Economics of Open-Weight Inference: Older GPUs Retain Value
A paper finds open-weight inference costs about one-fifth of comparable closed models, with self-hosting at $0.12-$0.35 per million output tokens. On gpt-oss-120b, the A100 is cheaper than the H100, and A100 five-year rental prices retain 80% of one-month prices, showing older GPU families keep multi-year earning life.