← Rynler blog

Integration guide ·

How to estimate LLM API costs: a token-first method

An AI assistant drafted these guides and checked them against public documentation. They were not generated through a recorded Rynler inference run and contain no performance or cost measurements.

Most LLM API cost estimates start from a rate card and a guess. Someone picks a model, assumes a typical request is "about a page of text", multiplies by expected traffic and writes a monthly figure into a planning document. The number looks precise, but almost every input to it was assumed rather than observed.

This guide describes a method that starts from tokens instead. You measure what one representative run actually consumes, keep input and output as separate quantities, model how your workload repeats that run, and only then apply the rates that are current when you make the decision. After launch, you compare the estimate with recorded usage and correct it.

Rynler publishes this method, and it was checked against public documentation. It contains no measured runs, prices or savings claims. The token figures in the worked example are illustrative assumptions, chosen to show the arithmetic.

Why LLM cost estimates go wrong

Estimates rarely fail because of a multiplication error. They fail because the quantity being multiplied does not describe the workload.

  • Input and output are averaged together. Many models charge different rates for input and output tokens, so a single blended "tokens per request" figure hides which side drives the bill.
  • Conversation history is resent. A chat API is stateless. Each turn sends the system instructions and the previous messages again, so input grows with every turn even when each user message is short.
  • Retrieved context is counted once instead of every call. In a retrieval workflow, the passages you add to the prompt are input on every request that includes them.
  • Agent tasks fan out. One user action can produce several model calls, each resending instructions, tool definitions and intermediate results.
  • Failures and retries are left out. A retried request is a new request. An attempt with an uncertain outcome may still have consumed work.

Each of these errors is a shape problem, not a pricing problem. The fix is to measure and model the shape before applying any rate.

Step 1: measure one representative run

Before estimating anything, send a small number of requests that look like real traffic: realistic instructions, a realistic amount of history or context, and the max_tokens value you intend to use. Then read the token counts the API reports, rather than counting characters or words yourself.

Rynler's chat completion response follows the OpenAI-compatible shape and includes a usage object with prompt_tokens, completion_tokens and total_tokens. When streaming, the final event carries usage before [DONE]. Record the two separate values for each run, not only the total.

{
  "usage": {
    "prompt_tokens": 1340,
    "completion_tokens": 248,
    "total_tokens": 1588
  }
}

The values above are an example of the shape, not a recorded response. Use the reported numbers because local counts are approximations: the request limits guide notes that token estimation is approximate and that the model's tokenizer determines actual usage. A local tokenizer is useful for checking whether a prompt fits a context window. It is not a substitute for the count the service bills against.

A single request is rarely representative. Measure a handful of runs that cover the short, typical and long cases your application will see, and keep the distribution rather than only the average. If one long-tail request type is ten times larger than the rest, it deserves its own line in the estimate.

Step 2: keep input and output as separate budgets

Write every estimate as two numbers: input tokens and output tokens. Rates are quoted separately for each, and the ratio between them differs from one workload to the next. A summarization job reads a lot and writes a little. A drafting tool can do the opposite. Merging them into one figure makes it impossible to apply the right rate later or to see which lever matters.

Output length is also partly under your control. The max_tokens parameter caps how many tokens a response may contain, and on Rynler it also bounds the amount held while a request runs: each request reserves its maximum possible cost and then settles against the tokens actually reported. A generous max_tokens therefore does not raise the final charge for a short answer, but it does raise the amount of prepaid credit temporarily held. Estimate output from the measured completion_tokens, and size the cap separately for safety and for credit headroom.

Some models also have a separate cache category for input. Rynler's model pages count input, cache and output separately, and the pricing page shows cache rates where they are supported. Do not assume a cache discount in a first estimate. Add it only once you know which part of your prompt is eligible and have recorded usage that shows the effect.

Step 3: model the workload shape

A measured run tells you what one call costs in tokens. The workload shape tells you how many calls a unit of user value needs and how their inputs grow. Pick the shape that matches your application and estimate per task, not per request.

WorkloadWhat repeatsWhat growsCalculator
Chat assistantOne call per user turnHistory resent on every turnChat API cost calculator
Retrieval (RAG)One call per questionRetrieved passages added as inputRAG cost calculator
AgentSeveral calls per taskInstructions, tool definitions and tool results resent per stepAgent cost calculator
Batch or pipelineOne call per document or itemUsually stable; output size varies by taskMonthly cost calculator

For chat, measure a typical conversation length in turns and remember that turn five sends turns one to four again. For retrieval, measure how many passages you add and how long they are; trimming context is often the largest lever on input. For agents, count the calls in a typical task, including planning, each tool step and the final answer. Tool definitions and history count toward input on every call that includes them, as the request limits guide states for the context-window check.

The calculators hub lists further variants by use case and model. Treat any calculator as a structure for your own measured numbers rather than as a source of them.

Step 4: multiply by cadence

Once you have tokens per task, multiply by how often the task runs. A simple monthly projection is runs per day multiplied by 30, applied separately to input and output.

Here is a worked example for a support chat, using assumed numbers:

  • The system instructions are 600 tokens.
  • A typical conversation has 6 turns. Each user message is about 80 tokens and each answer about 250 tokens.
  • Each turn resends the instructions and all previous messages, so turn 1 sends 680 input tokens and every later turn adds 330 more: 680, 1,010, 1,340, 1,670, 2,000 and 2,330.
  • Input per conversation is therefore 9,030 tokens. Output per conversation is 6 × 250 = 1,500 tokens.
  • At 2,000 conversations per day, a 30-day month has 60,000 conversations.
  • Monthly input: 60,000 × 9,030 = 541,800,000 tokens. Monthly output: 60,000 × 1,500 = 90,000,000 tokens.

Then multiply each total by its current rate for the model you selected, and add the two results. Keep the token totals in the planning document alongside the cost figure, because rates change and the token totals let anyone recompute the estimate.

Notice what history resend did to the example. If you had counted only the instructions and the new user message on each turn, you would have estimated 6 × 680 = 4,080 input tokens per conversation, less than half the modelled figure. For long conversations the gap grows further, which is why conversation length deserves an explicit cap or summarization strategy.

Run the same arithmetic for the high case as well as the typical one. A budget built only on the average is exceeded every time traffic skews toward long conversations or large documents.

Step 5: list what the formula leaves out

Tokens multiplied by rates describe successful work at the time you checked the rates. A budget also needs the items that formula does not cover.

  • Retries and failed requests. Each new request is separately billable. On Rynler, reusing an idempotency key that is already reserved returns IDEMPOTENCY_CONFLICT without another call or debit, while a new key starts a separate, potentially billable request. Budget for the retry rate you actually observe.
  • Uncertain outcomes. A timeout does not prove that nothing was consumed. Rynler keeps the reservation of an uncertain request until it is reviewed, so held credit can temporarily reduce the available balance.
  • Cache rules. Cache categories apply only where a model supports them and only to eligible input. Model them separately once observed.
  • Rate changes. Rynler prices are quotes that expire, so read the current rates on the models page when you finalize the estimate rather than reusing a figure from an older document.
  • Prepaid credit and minimum purchases. Rynler usage is prepaid, credit purchases are in USD, and the pricing page states the minimum credit purchase. Your cash outlay follows top-ups, not individual requests.
  • Taxes and payment fees. Check whether any apply to your purchases. They are not part of token usage and do not appear in a per-request estimate.

Writing these items down next to the token estimate keeps the planning figure honest. Some of them, such as retries, can be expressed as a percentage on top of the token totals once you have data.

Step 6: reconcile after launch

An estimate is a hypothesis. After launch, compare it with what the service recorded. Group the reported prompt_tokens and completion_tokens by task type, compare them with the per-task figures you modelled, and look for the shape errors listed at the start of this guide.

On Rynler, check the response usage, the workspace debit and any held reservation, as the migration guide recommends before moving traffic. Keep an estimate, recorded usage and a settled debit as three distinct quantities. When they disagree, the cause is usually a different context size, a longer conversation, more agent steps or additional retries than you assumed.

Uncertain requests need particular care in this comparison. The guide on reconciling usage after LLM API timeouts explains how to keep delivery, recorded usage, held credit and settled charges apart so that unresolved attempts neither disappear from the report nor get counted as final charges.

Update the estimate with what you learn. Replace assumed token counts with observed distributions, and record the date of the rates you applied.

Checklist

  • Measure several representative runs and record prompt_tokens and completion_tokens separately.
  • Keep input and output as two budgets throughout the estimate.
  • Choose the workload shape and estimate per task, including history resend, retrieved context and agent steps.
  • Multiply by runs per day and by 30, for both the typical and the high case.
  • Apply the current rate for each token type only at the end, and record when you read it.
  • List retries, uncertain outcomes, cache treatment, minimum purchases, taxes and fees next to the token totals.
  • Set max_tokens deliberately, knowing it affects the amount held while a request runs.
  • After launch, compare the estimate with recorded usage and debits, and keep unresolved requests visible.

Next steps

Read the current rates on the models page and the reservation and credit terms on the pricing page. To measure your first representative run, follow the quickstart, which has you create the API key with a spending limit, and read the reported usage from the response. A cost estimate you can recompute from token totals is more useful than a single monthly number that nobody can trace.