Comparison · Updated
The real cost of an LLM API in 2026: fees, minimums, refunds and discounts
The per-token price is the number every comparison table shows. It is rarely the whole bill. Purchase fees, minimum top-ups, expiring credits, refund rules, cache and batch discounts, time-of-day pricing and cost reservations all change what you actually pay. This page lists those rules for ten APIs from their own documentation, read on 2026-10-07.
Rynler publishes this page and is one of the ten. Its own rules are in the table, including the ones that are not in its favor. No per-token price appears here: they change too often to freeze in a page. Compare live prices on each provider's site.
The rules that change the bill
| API | Charge on top of model prices | Minimum purchase | Credit expiry | Refund of unused credit | Cached input | Batch discount |
|---|---|---|---|---|---|---|
| OpenRouter | 5.5% card fee on credits, USD 0.80 minimum; 5% crypto | Not checked | May expire after one year | Card purchases: requests within 24 hours, fees not refunded; crypto never | Provider rules passed through | Typically 50%, 24-hour window |
| Requesty | 5% | No minimum commitment; purchase minimum not published | Not published for purchased credits | Final, non-refundable | Up to 90% savings on hits with auto_cache | Not published |
| Eden AI | 5.5% platform fee at checkout | Not published | Not published | Non-refundable unless otherwise stated | Provider caching, no rate published | Not published for LLMs |
| Hugging Face | No markup, no additional fees | Not published | Not published | Not published | Not published as its own feature | Not published |
| OpenAI | No fee published | USD 5 | One year | Non-refundable except listed cases | On by default, up to 95% off | 50%, 24-hour window |
| Together AI | No fee published | USD 5 | Not currently | Fees paid non-refundable (terms) | Automatic on models with a cached rate | Up to 50%, two listed models |
| Groq | No fee published | Postpaid, invoiced in arrears | Not applicable | Case by case | 50% on three models | 50%, 24 hours to 7 days |
| Fireworks AI | No fee published | Not published | Not published | Fees paid non-refundable (terms) | On by default, 50% default discount | 50% |
| DeepSeek | No fee published | Not published | Top-up balance does not expire | Unused balance refundable on request, subject to review, after handling fees | Automatic, separate cache-hit rate | Not published |
| Rynler | No purchase fee | Credits from USD 10, per the pricing page | Never | Non-refundable except where the law requires | Separate cache-read rate on supported models | No batch endpoint |
"Not published" means we did not find it on the official pages we read, not that the policy does not exist. "Not checked" means we did not look for it. "No fee published" is the same caution applied to fees.
1. Fees on buying credits
Most routers either mark up usage or charge when you buy credits; Hugging Face states it does neither. The effect is the same: a percentage on everything you spend. On USD 300 of monthly usage, a 5.5% fee is USD 16.50 a month; a 5% fee is USD 15. Direct providers and Hugging Face did not publish such a fee on the pages we read, and Rynler credits the exact amount you pay. Taxes and your bank's foreign-exchange fees are separate everywhere.
Requesty describes its 5% in two ways: a markup on usage on its pricing page, and a service fee added to each top-up in its billing documentation. The cost to you is about the same either way.
2. Minimums and expiry
A minimum purchase matters when you only want to test. OpenAI and Together publish USD 5; Rynler starts at USD 10. Expiry matters more over time: OpenAI's purchased credits expire after one year, and OpenRouter reserves the right to expire unused credits after one year. Together says its prepaid credits do not currently expire, DeepSeek says top-up balances do not expire, and Rynler credits do not expire.
3. Refunds
Most providers treat credit purchases as final. The exceptions we found: OpenRouter accepts refund requests within 24 hours for card purchases (platform fees excluded, crypto never refundable), and DeepSeek reviews refund requests for the unused balance and, if approved, refunds it after deducting handling fees, with no partial refunds (terms). Rynler is in the majority: credits are non-refundable except where the law requires a remedy. If you are unsure about a provider, buy the minimum first.
4. Caching: when the cheapest rate is not the input rate
Cached input can cost far less than fresh input. OpenAI advertises discounts of up to 95% with caching on by default (guide), Fireworks a default 50% (guide), Groq 50% on three models (guide). Two consequences for your bill:
- Prompt order matters. Caches match prefixes: put the stable part of the prompt (instructions, documents) first and the changing part last.
- Hits are not guaranteed. DeepSeek and Groq both say a cache hit is best effort, and cache entries expire after minutes or hours. Budget with the uncached rate and treat cache savings as upside.
Rynler lists a separate cache-read rate for models where it is supported, on the live pricing page. A measured example is planned for our blog; we will not quote a saving before we have run it.
5. Batch discounts
OpenAI, Groq and Fireworks publish 50% off for asynchronous batch jobs, OpenRouter says batch requests are typically billed at 50% (batch quickstart), and Together publishes up to 50% on two listed models. The price is latency: results within 24 hours (OpenAI, Together) or up to seven days (Groq). Groq also says its batch discount does not stack with caching. Rynler has no batch endpoint, so a nightly job costs the same as a real-time request, except on the standard DeepSeek models it lists, whose off-peak rates are lower (section 6).
6. Time-of-day pricing
DeepSeek charges off-peak rates at half its peak rates. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays (pricing). Scheduling large jobs outside those windows halves that part of the bill. On Rynler, the standard DeepSeek models it lists (DeepSeek V4 Flash and DeepSeek V4 Pro) follow the official peak and off-peak schedule; DeepSeek V4.1 Flash is a discounted model with a flat price. None of them could be called on 2026-10-07.
7. Reservations and held money
Prepaid APIs need a way to stop you from spending more than your balance. Rynler reserves the maximum possible cost of each request before calling the model, then settles against the tokens actually reported. Two effects you should plan for:
- Peak balance matters. Many concurrent requests with high
max_tokensreserve more credit than they will spend. Setmax_tokensclose to what you need. - Uncertain requests hold money. If a response or its usage cannot be verified, the reservation stays held until reviewed, rather than being charged twice (request limits).
Postpaid billing, like Groq's monthly invoice, avoids holds but moves the risk to your card, charged at month end or, for new accounts, at progressive usage thresholds.
A checklist before you pick an API on price
- Multiply your monthly spend by the purchase fee or markup, if any.
- Check the minimum purchase and whether credits expire before you will use them.
- Read the refund clause; assume non-refundable unless stated otherwise.
- Reorder prompts for caching, then budget without it.
- Move delay-tolerant work to a batch endpoint or off-peak hours where available.
- Set
max_tokensand spending limits to keep reservations and surprises small.
Sources and method
External sources, read on 2026-10-07 and subject to change, are linked in the table and next to each statement. Some terms pages only render in a browser; we quoted them as rendered. Rynler facts come from our pricing, request limits and terms. The fee examples are arithmetic on published percentages, not invoices. Corrections: contact@rynler.com.
Estimate the cost
- All LLM cost calculators
Calculators by provider, workload and budget, all built on the same two-rate token formula.
- How to estimate LLM API costs
The token-first method: measure one run, model the workload shape, project monthly, reconcile usage.