← All comparisons

Comparison · Updated

The real cost of an LLM API in 2026: fees, minimums, refunds and discounts

The per-token price is the number every comparison table shows. It is rarely the whole bill. Purchase fees, minimum top-ups, expiring credits, refund rules, cache and batch discounts, time-of-day pricing and cost reservations all change what you actually pay. This page lists those rules for ten APIs from their own documentation, read on 2026-10-07.

Rynler publishes this page and is one of the ten. Its own rules are in the table, including the ones that are not in its favor. No per-token price appears here: they change too often to freeze in a page. Compare live prices on each provider's site.

The rules that change the bill

APICharge on top of model pricesMinimum purchaseCredit expiryRefund of unused creditCached inputBatch discount
OpenRouter5.5% card fee on credits, USD 0.80 minimum; 5% cryptoNot checkedMay expire after one yearCard purchases: requests within 24 hours, fees not refunded; crypto neverProvider rules passed throughTypically 50%, 24-hour window
Requesty5%No minimum commitment; purchase minimum not publishedNot published for purchased creditsFinal, non-refundableUp to 90% savings on hits with auto_cacheNot published
Eden AI5.5% platform fee at checkoutNot publishedNot publishedNon-refundable unless otherwise statedProvider caching, no rate publishedNot published for LLMs
Hugging FaceNo markup, no additional feesNot publishedNot publishedNot publishedNot published as its own featureNot published
OpenAINo fee publishedUSD 5One yearNon-refundable except listed casesOn by default, up to 95% off50%, 24-hour window
Together AINo fee publishedUSD 5Not currentlyFees paid non-refundable (terms)Automatic on models with a cached rateUp to 50%, two listed models
GroqNo fee publishedPostpaid, invoiced in arrearsNot applicableCase by case50% on three models50%, 24 hours to 7 days
Fireworks AINo fee publishedNot publishedNot publishedFees paid non-refundable (terms)On by default, 50% default discount50%
DeepSeekNo fee publishedNot publishedTop-up balance does not expireUnused balance refundable on request, subject to review, after handling feesAutomatic, separate cache-hit rateNot published
RynlerNo purchase feeCredits from USD 10, per the pricing pageNeverNon-refundable except where the law requiresSeparate cache-read rate on supported modelsNo batch endpoint

"Not published" means we did not find it on the official pages we read, not that the policy does not exist. "Not checked" means we did not look for it. "No fee published" is the same caution applied to fees.

1. Fees on buying credits

Most routers either mark up usage or charge when you buy credits; Hugging Face states it does neither. The effect is the same: a percentage on everything you spend. On USD 300 of monthly usage, a 5.5% fee is USD 16.50 a month; a 5% fee is USD 15. Direct providers and Hugging Face did not publish such a fee on the pages we read, and Rynler credits the exact amount you pay. Taxes and your bank's foreign-exchange fees are separate everywhere.

Requesty describes its 5% in two ways: a markup on usage on its pricing page, and a service fee added to each top-up in its billing documentation. The cost to you is about the same either way.

2. Minimums and expiry

A minimum purchase matters when you only want to test. OpenAI and Together publish USD 5; Rynler starts at USD 10. Expiry matters more over time: OpenAI's purchased credits expire after one year, and OpenRouter reserves the right to expire unused credits after one year. Together says its prepaid credits do not currently expire, DeepSeek says top-up balances do not expire, and Rynler credits do not expire.

3. Refunds

Most providers treat credit purchases as final. The exceptions we found: OpenRouter accepts refund requests within 24 hours for card purchases (platform fees excluded, crypto never refundable), and DeepSeek reviews refund requests for the unused balance and, if approved, refunds it after deducting handling fees, with no partial refunds (terms). Rynler is in the majority: credits are non-refundable except where the law requires a remedy. If you are unsure about a provider, buy the minimum first.

4. Caching: when the cheapest rate is not the input rate

Cached input can cost far less than fresh input. OpenAI advertises discounts of up to 95% with caching on by default (guide), Fireworks a default 50% (guide), Groq 50% on three models (guide). Two consequences for your bill:

  • Prompt order matters. Caches match prefixes: put the stable part of the prompt (instructions, documents) first and the changing part last.
  • Hits are not guaranteed. DeepSeek and Groq both say a cache hit is best effort, and cache entries expire after minutes or hours. Budget with the uncached rate and treat cache savings as upside.

Rynler lists a separate cache-read rate for models where it is supported, on the live pricing page. A measured example is planned for our blog; we will not quote a saving before we have run it.

5. Batch discounts

OpenAI, Groq and Fireworks publish 50% off for asynchronous batch jobs, OpenRouter says batch requests are typically billed at 50% (batch quickstart), and Together publishes up to 50% on two listed models. The price is latency: results within 24 hours (OpenAI, Together) or up to seven days (Groq). Groq also says its batch discount does not stack with caching. Rynler has no batch endpoint, so a nightly job costs the same as a real-time request, except on the standard DeepSeek models it lists, whose off-peak rates are lower (section 6).

6. Time-of-day pricing

DeepSeek charges off-peak rates at half its peak rates. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays (pricing). Scheduling large jobs outside those windows halves that part of the bill. On Rynler, the standard DeepSeek models it lists (DeepSeek V4 Flash and DeepSeek V4 Pro) follow the official peak and off-peak schedule; DeepSeek V4.1 Flash is a discounted model with a flat price. None of them could be called on 2026-10-07.

7. Reservations and held money

Prepaid APIs need a way to stop you from spending more than your balance. Rynler reserves the maximum possible cost of each request before calling the model, then settles against the tokens actually reported. Two effects you should plan for:

  • Peak balance matters. Many concurrent requests with high max_tokens reserve more credit than they will spend. Set max_tokens close to what you need.
  • Uncertain requests hold money. If a response or its usage cannot be verified, the reservation stays held until reviewed, rather than being charged twice (request limits).

Postpaid billing, like Groq's monthly invoice, avoids holds but moves the risk to your card, charged at month end or, for new accounts, at progressive usage thresholds.

A checklist before you pick an API on price

  • Multiply your monthly spend by the purchase fee or markup, if any.
  • Check the minimum purchase and whether credits expire before you will use them.
  • Read the refund clause; assume non-refundable unless stated otherwise.
  • Reorder prompts for caching, then budget without it.
  • Move delay-tolerant work to a batch endpoint or off-peak hours where available.
  • Set max_tokens and spending limits to keep reservations and surprises small.

Sources and method

External sources, read on 2026-10-07 and subject to change, are linked in the table and next to each statement. Some terms pages only render in a browser; we quoted them as rendered. Rynler facts come from our pricing, request limits and terms. The fee examples are arithmetic on published percentages, not invoices. Corrections: contact@rynler.com.

Estimate the cost