Illustrative model
AI calculator
Model indicative annual token cost under a raw frontier path versus a governed path with routing and caching. Purely client-side — nothing is collected or transmitted.
Your inputs
Consumption parameters
Adjust the sliders to reflect a representative workload. Outputs update live.
Completed requests per month for one workload or estate slice.
Combined prompt and completion tokens, indicative average.
Blended indicative rate for the unoptimised path.
Portion of queries that can use a lower-cost model without breaking quality gates.
Indicative rate for the routed / fine-tuned path.
Share of queries served from cache without a full model call.
Based on your inputs
Indicative annual difference
$0
0% reduction
Based on your inputs. Illustrative model only.
This calculator provides an illustrative model based on your inputs. Actual savings depend on model selection, query complexity, caching effectiveness, architecture and operational maturity. It is not a quote or guarantee.
Method
How the numbers are derived
- Unoptimised annualMonthly queries × tokens × frontier $/1M tokens, annualised.
- Governed annualCache reduces billed queries; remaining volume is split between frontier and smaller-model rates by the routing share.
- Cost per outcomeAnnual cost divided by annual query volume — a unit view for leadership conversations.
Honest limits
What this model is not
- Not a quote, proposal or guaranteed saving
- Does not include infrastructure, retrieval or human-review cost
- Does not validate quality after routing or caching
- Does not replace metered gateway telemetry from your estate
For a scoped FinOps assessment, start with a readiness discussion or the self-assessment on AI Value.