Rate limits
The API is rate limited to keep it fast and fair for everyone. Limits are applied using a leaky bucket algorithm and are aggregated per restaurant (across all of that restaurant’s API keys).
The limits
Section titled “The limits”Reads and writes draw on two separate budgets, and a request spends exactly
one of them. GET, HEAD and OPTIONS spend the read budget; every other
method — POST, PATCH, DELETE — spends the write budget. A write never also
spends read quota.
| Limit | Reads | Writes |
|---|---|---|
| Sustained rate | ~2 requests per second | ~1 request per second |
| Burst | up to 60 requests | up to 20 requests |
| Daily cap | ~25,000 requests per day | ~2,000 requests per day |
The leaky bucket lets you briefly burst, then drains at the sustained rate — so steady traffic flows smoothly, while spikes are smoothed out rather than hard-blocked at the first extra request.
Writes are the lower budget because a create or a modify takes a lock on the restaurant’s service day while it searches for a table. Everything else writing that day — the back office, the widget, the restaurant’s own staff — queues behind it. A queue of 250 writes replayed after an outage still drains in about four minutes.
Rate limit headers
Section titled “Rate limit headers”When a response reports the limit state, it does so in four headers:
RateLimit-Limit: 60RateLimit-Remaining: 42RateLimit-Reset: 9RateLimit-Resource: read| Header | Meaning |
|---|---|
RateLimit-Limit |
The bucket capacity (burst size). |
RateLimit-Remaining |
Requests you can still make right now. |
RateLimit-Reset |
Seconds until capacity is fully replenished. |
RateLimit-Resource |
Which budget the other three describe: read or write. |
The triple always describes the budget this request spent. Pace reads off a
read response and writes off a write response; reading RateLimit-Remaining
without checking RateLimit-Resource means watching two unrelated series
through one name.
A response can arrive with none of these headers. When RateLimit-Remaining is
absent, read it as “not reported”, not as “unlimited”: pace to the limits in the
table above, and keep honouring 429 and Retry-After.
When you are throttled
Section titled “When you are throttled”Exceeding the limit returns 429 Too Many Requests with a Retry-After
header (seconds to wait) and a rate_limit_error envelope:
HTTP/1.1 429 Too Many RequestsRetry-After: 1RateLimit-Limit: 20RateLimit-Remaining: 0RateLimit-Reset: 1RateLimit-Resource: writeRateLimit-Resource names the budget you exhausted. It reads ip when neither
per-restaurant bucket refused the call and a per-address limit did, which is a
different problem from your key’s quota.
{ "error": { "type": "rate_limit_error", "code": "rate_limited", "message": "Too many requests. Retry after 1 second.", "doc_url": "https://docs.useservice.app/api/errors#rate_limited" }}Staying within the limits
Section titled “Staying within the limits”- Honour
Retry-After. On a429, wait the indicated number of seconds before retrying. - Back off exponentially on repeated
429s, with a little jitter. - Watch
RateLimit-Remaining, alongsideRateLimit-Resource, and slow down before you hit zero. - Pace writes separately. A loop that creates bookings as fast as it reads them exhausts the write budget first, and a retry storm reaches the write burst in under a second.
- Sync incrementally. Use
updated_sinceinstead of repeatedly re-reading full collections, and prefer webhooks over tight polling loops.