Browse any relay's pricing page and you'll run into "1.5x multiplier," "group rate," or "20% off metered usage." The billing multiplier is the core mechanic behind how relays price things, and it's also the number newcomers get wrong most often. Here's the plain-language version.
What the multiplier actually is
The multiplier is the markup a relay applies on top of official pricing. The formula is simple:
What you pay = official API price × multiplier
Example: say a model's official price is $3 per million input tokens. At a 2x multiplier, the same usage costs you $6. At 1.2x, it's $3.60. The lower the multiplier, the closer you are to official pricing.
So why isn't "lower is always better" the whole story? Because a low multiplier means thin margins for the relay — which means either they're making it up on volume, or they're making it up somewhere else. That "somewhere else" is the next section.
Three common ways low multipliers get made up elsewhere
1. Input and output priced separately
Official pricing charges different rates for input and output tokens (output is usually several times pricier). Some providers advertise "1.2x on input" front and center while quietly listing output at 3x. In real conversations, output tokens are the bulk of the cost — so the actual bill isn't nearly as cheap as the headline number suggests.
2. Different models or groups, different multipliers
The "20% off" banner on the homepage is often the loss-leader rate for an obscure or older model. The flagship model you actually want — the latest Claude or GPT — often lives in a separate "premium group" with a much higher multiplier.
3. Silent model substitution
The worst version: you request the flagship model, but the backend quietly routes you to a smaller, cheaper, or distilled model instead, and quality drops noticeably. The multiplier is low, but you're not getting what you thought you paid for.
How to work out the real cost
Don't take the homepage number at face value — check it yourself in three steps:
- Look up the multiplier for the exact model you'll use, not the loss-leader rate on the landing page;
- Check input and output multipliers separately, and weight them by your actual usage pattern (chat-heavy workloads skew output, extraction-heavy ones skew input);
- Run a small top-up against a real workload, then divide what you were actually billed by the official price to back into the effective multiplier — that number is the one that counts.
Beyond the multiplier: don't forget the hidden costs
A 0.1x cheaper multiplier isn't much of a deal if the service drops every other day, forces retries, and nobody's around at 2am. That time and reliability cost usually outweighs the price difference by a wide margin. For production use, stability generally beats a rock-bottom price. A provider at 1.5x that just works is often cheaper in practice than one at 1.1x that constantly misbehaves.
The short version
Multiplier = official price × markup factor, but the number that actually matters is model-specific, split between input and output, and confirmed against your real bill. Once you understand how it works, you can spot most low-price traps on sight. Weigh it alongside stability and how safe your deposit is — see 7 things to check before choosing a relay, or head to the HowToken directory to compare providers by minimum top-up and payment method.