Every paid data API publishes a rate ceiling, and most buyers skim past it. That is a mistake, because the ceiling silently decides what you can build. A ten-second polling loop across forty instruments is a different product from a one-minute loop across five, and the ceiling is what separates them.

Read the ceiling as a budget

Work out your worst case before you subscribe, not after. Instruments multiplied by polls an hour, plus whatever your retries add during an outage. Retries are the part people forget: a failing upstream is exactly when your client hammers hardest, so the ceiling binds worst at the moment you most need data.

A ceiling quoted per hour rather than per minute is more forgiving of bursts, because a short spike is absorbed by the rest of the window. A per-second cap is stricter than a per-hour cap that works out to the same total.

What should happen when you hit it

A refusal should be machine-readable. An HTTP 429, a named reason code your code can branch on, and a Retry-After header telling you when to come back. Prose is not an interface: a client cannot parse "you have made too many requests lately" and decide anything sensible.

The refusal should also never be cached for somebody else. A rate-limit answer is about one caller, and a shared cache holding it turns your problem into everyone’s.

Bucketed by key, not by address

Where the limit is counted matters as much as its size. Counting by IP address means everyone behind the same office network or cloud egress shares a ceiling. Counting by API key means the ceiling you bought is the ceiling you get, which is the entire point of paying for one.

If a vendor cannot tell you which it uses, assume the cheaper one.