01 Workload lifecycle & sizing / TCO
- LLM Inference Benchmarking: How Much Does Your Inference Cost?NVIDIA Technical Blog
Turn latency/throughput curves into a cost-per-token TCO model with GenAI-Perf.
- Rethinking AI TCO: Why Cost per Token Is the Only MetricNVIDIA Blog
The right TCO unit for inference fleets, and the framework to anchor sizing to.
- ZeRO: Memory Optimizations Toward Trillion-Parameter ModelsRajbhandari et al. (arXiv)
Canonical primary source for training memory math (params + grads + optimizer + activations).