Economics
Latency and cost are the same trade
Bigger batches raise tokens per second and raise time-to-first-token with them. Aggressive quantisation cuts memory and can cut quality. We agree the budgets first, in milliseconds and in currency, then tune against them.
- TTFT and tail latency budgets agreed before tuning
- Cost tracked per million tokens and per conversation
- Every change gated on your eval set, not a benchmark