freeseek

Cheap tokens are the point

This service resells nothing and marks up nothing, so its existence hangs on a single number: what a token costs. That number has quietly become one of the widest price spreads in software. As of August 2026, a million output tokens from deepseek-v4-flash cost $0.28. The same million from a frontier flagship cost $25 to $50 — $180 from the pro tiers. The models are not equivalent, but on everyday work they are far closer in capability than they are in price.

The rate cards

Published list prices, USD per million tokens, checked against each vendor's own pricing page on 2026-08-06:

modelinputcached inputoutput
deepseek-v4-flash$0.14$0.0028$0.28
deepseek-v4-pro$0.435$0.003625$0.87
GPT-5.2$1.75$0.175$14.00
Claude Sonnet 5*$2.00$0.20$10.00
GPT-5.4$2.50$0.25$15.00
Claude Opus 5$5.00$0.50$25.00
GPT-5.5$5.00$0.50$30.00
Claude Fable 5$10.00$1.00$50.00
GPT-5.5-pro$30.00$180.00

*Introductory pricing through 2026-08-31; $3 / $15 after. Sources: DeepSeek, Anthropic, OpenAI. Prices change monthly; when this table and a vendor's page disagree, the vendor's page is right.

Read down the output column: flash to Fable 5 is 179×. The cached-input column is starker — $1.00 against $0.0028 is 357× — and DeepSeek's cache is automatic and free to write, where Anthropic bills cache writes at 1.25–2× the input rate.

The same task

Rate cards mislead without a workload, so take an ordinary agent exchange — the shape Anthropic itself uses as a worked example: 50k input tokens of which 40k are cache reads, 15k output tokens. Counting the same nominal tokens at each vendor's list prices:

modelthat exchange costsvs flash
deepseek-v4-flash$0.0057
Claude Sonnet 5$0.1831×
GPT-5.4$0.2646×
Claude Opus 5$0.4578×
GPT-5.5$0.5291×
Claude Fable 5$0.89156×
GPT-5.5-pro$4.20735×

Not twenty percent cheaper. Thirty to a few hundred times cheaper, depending on the model and how much of the prompt caches.

What this does not claim

Per-token is not per-task. A stronger model that solves a hard problem in one attempt can beat a cheaper model that needs three, and on the hardest work the frontier models earn their price. Tokenizers differ too — Anthropic notes that its current tokenizer produces roughly 30% more tokens for the same text than its previous one, so identical work is not identical token counts across vendors. And DeepSeek has announced peak-hour pricing at 2× the listed rates, effective date not yet set. The honest claim is narrower and still remarkable: for the broad middle of real work — summarize, translate, refactor, answer, glue — the going rate differs by two orders of magnitude depending on whose API you call.

Why it matters

Chat is measured in thousands of tokens; agents are measured in millions. The moment a model works unattended — reading files, retrying, checking its own output — token consumption stops tracking human attention and starts tracking machine patience. An overnight agent run that emits ten million output tokens costs $2.80 at flash prices and $500 at Fable prices. One of those is "leave it running"; the other is a line item that gets a meeting. At frontier prices, autonomy is a luxury good. At flash prices, it is a background process.

Whatever AGI turns out to be, it will be made of tokens, and nobody runs civilization-scale inference at $50 per million. Every 10× drop in token price makes a class of applications viable that was silly the day before — the same way compute-per-dollar curves, not any single breakthrough, decided what software got built. Cheap tokens are not the budget option. They are the substrate.

This page is also the explanation of the gateway you are reading it on. At flash prices, a dollar buys roughly three thousand ordinary conversational turns; at frontier list prices, the same dollar buys about fifty. A free tier funded by donated keys and pocket money is arithmetic that only works at the bottom of that table — which is why it runs on deepseek-v4-flash, and why there is no paid tier to upsell you to.