Rate cards are only the first layer
A headline input-token price does not describe the full cost of a workload. Output tokens may carry a different rate, cached inputs can be discounted, long contexts may use separate tiers, and image or video requests often depend on resolution or duration.
Compare the billable units your application actually produces rather than a single published number.
Build a workload-shaped comparison
Use a sample of real or representative requests and record:
- Average input, cached input, and output volume.
- Success rate and the number of retries per completed task.
- Latency to first token and total response time.
- Model quality on the task's acceptance criteria.
- Any gateway, storage, or data-transfer charges.
Normalize to a useful business unit
Convert raw token spend into cost per completed task, generated asset, support resolution, or active user. This gives product and finance teams a metric they can reason about and prevents token-level savings from hiding lower completion rates.
Account for changing market conditions
A one-time spreadsheet becomes stale as models and supply change. Keep pricing, quality, and availability measurements separate, then let routing policies balance them for each workload. DIT's exchange approach is designed to evaluate qualified supply paths dynamically instead of locking every request to one fixed rate card.
Use one DIT key to access supported models through a market of qualified AI providers.
Get your API key
