The idea
AI services may bill for input tokens, output tokens, requests or other units. Identify the actual units and rate period before estimating usage. Include retrieved context, retries and supporting services. Prices vary by provider and model, so use a dated official pricing source for a real budget, not an undated example.
Worked example
At hypothetical rates of 2 units per million input tokens and 8 per million output tokens, a request using 1,000 input and 500 output tokens costs 0.002 plus 0.004, or 0.006 units. This is an arithmetic example, not a quote for any live service.
Try it
Estimate 10,000 requests using the hypothetical request above. Then add a 20% retry overhead under the assumption that retries have equal usage. List other costs excluded from your estimate and state which rate information you would verify before a real deployment.
