Reference
Rate limits
Two per-minute ceilings, counted in the database rather than in a process, applied to every /api/v1 route by the same gate that checks the key: one for the API key, and one for the tenant behind all of them.
Status
Every /api/v1 request that carries a credential passes through one gate, which authenticates the key, checks the route's scope, meters the call and resolves the tenant's mode. There is no route where a keyed client can skip the budget.
Two routes are outside the key gate by design, and both are browser surfaces rather than API calls: GET /api/v1/images/{id} is an <img> tag, and GET /api/v1/stream/prices is an EventSource. Neither can hold a key. Neither is unbounded either. Images are metered per client address at 300 a minute, which is several full catalogue pages inside a minute, and the bytes are cached so a burst does not become one upstream fetch per request. Over that, images answer 429 rate_limited like anything else. The stream caps connections, contracts per connection and total instruments, meters a session that presents one, and still applies an operator's offering policy. A stream at capacity answers 503 stream_capacity with a Retry-After.
| Endpoint | Scope |
|---|---|
GET /api/v1/categories | markets:read |
POST /api/v1/embed-sessions | orders:write |
GET /api/v1/events | markets:read |
GET /api/v1/events/{id} | markets:read |
POST /api/v1/funding/deposit-address | funding:write (sandbox only) |
GET /api/v1/funding/deposits/{id} | funding:write (sandbox only) |
GET /api/v1/funding/supported-assets | funding:write (sandbox only) |
GET /api/v1/images/{id} | No key and no scope. Metered by client address, not by tier. See below. |
GET /api/v1/orders | orders:read |
POST /api/v1/orders | orders:write |
GET /api/v1/orders/{id} | orders:read |
POST /api/v1/orders/{id}/cancel | orders:write |
POST /api/v1/orders/{id}/confirm-reserve | orders:write |
GET /api/v1/positions | positions:read |
GET /api/v1/positions/{id} | positions:read |
POST /api/v1/quotes | quotes:write |
GET /api/v1/reconciliation | ledger:read |
GET /api/v1/reconciliation/feed | ledger:read |
POST /api/v1/sandbox/resolutions | settlements:read (sandbox only) |
GET /api/v1/sandbox/scenarios | orders:read (sandbox only) |
POST /api/v1/sandbox/scenarios | orders:write (sandbox only) |
GET /api/v1/settlements | settlements:read |
GET /api/v1/settlements/positions | settlements:read (deprecated) |
GET /api/v1/stream/prices | No key and no scope. Metered by client address, not by tier. See below. |
GET /api/v1/users/{id}/balance | ledger:read (sandbox only) |
GET /api/v1/users/{id}/jurisdiction | markets:read |
PUT /api/v1/users/{id}/jurisdiction | orders:write |
GET /api/v1/users/{id}/ledger | ledger:read (sandbox only) |
DELETE /api/v1/users/{id}/limits | orders:write |
GET /api/v1/users/{id}/limits | orders:write |
PUT /api/v1/users/{id}/limits | orders:write |
That is the whole /api/v1 surface. This table is generated from the OpenAPI document at build time, and CI asserts that every scope in that document is the scope the handler actually passes to the gate, so it cannot list a route that does not serve, omit one that does, or name a scope that is not enforced. The version of this table that was maintained by hand had gone six routes out of date underneath a sentence claiming to be exhaustive.
How the limit works
| Property | Value |
|---|---|
| Scope | Two counters. One per API key, and one per operator across every key it holds. Neither is per IP. |
| Window | Fixed, 60 seconds, aligned to the wall clock minute. |
| Per-key default | 120 requests per minute on a sandbox key, 600 on a live one. |
| Per-key ceiling | A key cannot be configured above 6,000 per minute, by anyone. |
| Per-tenant ceiling | 12,000 requests per minute across every key the operator holds, unless a deployment sets a lower one. |
| Configured on | The key itself, so a sandbox key and a live key can differ. The tenant ceiling is not yours to set. |
| Counter | A row in Postgres, incremented by a single upsert that returns the new count. |
The second ceiling, and why your headers cannot see it
Minting more keys divides the tenant budget. It does not multiply it.
The per-key counter is keyed on the key. Nothing caps how many keys you mint, so N keys would have bought N times the allowance. A second counter, keyed on the operator, closes that: every key you hold draws on one tenant-wide budget.
The three RateLimit-* headers describe the key, deliberately and permanently: they are the numbers integrations already build backoff against, and repointing them at a different counter would change the meaning of a value in use. So a response can carry RateLimit-Remaining: 500 and the very next request can still be refused, because the tenant counter, which no header reports, is the one that ran out.
Both refusals are 429 rate_limited with retryAfterSeconds and a Retry-After header, so the correct action is the same either way: sleep that long and come back. The message is the only thing that differs and it is prose, so do not branch on it. If you need to know which ceiling you are hitting, sum your own request rate across every key rather than trying to read it off a response.
Scope is checked before the limit is consumed
A caller with the wrong scope has made a mistake that no amount of waiting fixes. Charging them a request to say so would turn the 403 into a 429 on the retry, which reads as an entirely different bug and sends whoever is debugging it to this page instead of to their key configuration.
Why the count lives in the database
A per-process counter multiplies the real limit by however many instances are running, which is a limit in name only. A single upsert returning the new count also means two concurrent requests cannot both read the same number and both decide they were under.
A fixed window admits a burst
Twice the allowance can pass across a window boundary. That is a known and accepted property: this protects the service from a runaway loop, not from an adversary. Do not design around the burst. It is not a feature and the window shape may change.
What every response tells you
Most responses carry the current window on three headers, a success as much as a refusal, so a client can slow down before it is refused rather than discovering the budget by hitting it. The exceptions are listed below.
RateLimit-Limit: 120
RateLimit-Remaining: 96
RateLimit-Reset: 37When the allowance is spent the request is refused with 429, a Retry-After header, and a body that repeats the wait in seconds:
{
"error": "rate_limited",
"message": "Over 120 requests per minute for this key.",
"retryAfterSeconds": 37
}RateLimit-Limit: the allowance for this key, this window.RateLimit-Remaining: floored at zero, so it never goes negative on a burst.RateLimit-Reset,Retry-AfterandretryAfterSeconds: seconds until the window rolls. Sleep that long, not a fixed guess.- Read the headers on successful responses too. Backing off as
RateLimit-Remainingapproaches zero is how you avoid most of the429s. It cannot avoid all of them: the headers report the key, the tenant ceiling is counted separately, and two money-moving routes answer without the headers at all. Keep a retry path. POST /api/v1/quotesandPOST /api/v1/ordersdo not return the three headers. Both are metered exactly like every other route; they just do not report the window back. Carry your own count between the reads that do.
The SDK raises a RateLimitError carrying retryAfterSeconds, so the wait is already parsed for you.
Staying inside the budget
Most integrations that hit a limit are doing one of four things.
| Instead of | Do this |
|---|---|
| Refetching the whole catalogue on a timer | Page it once with ?cursor=, then re-read narrowly with ?closesAfter= and ?tradable=true, rather than pulling every event to find the handful that changed. |
| Polling the settlement register continuously | Poll on the order of once a minute, carrying nextCursor. Settlements are not high-frequency events. |
| Re-quoting on a timer to keep a ticket warm | Quote when the player is about to act. A quote carries expiresAt; re-quote on expiry, not on an interval. |
| Polling every open order for a fill | The order response already carries the fill. The only order worth polling is one whose settlementState is pending, and the reconciliation feed reports every money fact in one resumable stream. |
Retry with jitter
If every client in a fleet sleeps for exactly retryAfterSeconds, they all return at the same instant and the next window is exhausted immediately. Add a random fraction of a second.

