Rate limits and quota¶
Every authenticated response tells you where you stand. Read the headers rather than counting requests yourself:
curl -sD - -o /dev/null https://api.trilocore.ai/api/v1/workspaces \
-H "Authorization: Bearer $TRILOCORE_API_KEY" \
| grep -i '^x-ratelimit\|^retry-after\|^x-request-id'
x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 597
x-ratelimit-reset-requests: 41
x-request-id: req_e345361146654a96898ff450881779d0
The headers¶
| Header | Meaning |
|---|---|
x-ratelimit-limit-requests |
requests allowed in the current window |
x-ratelimit-remaining-requests |
requests left in the current window |
x-ratelimit-reset-requests |
seconds until the window resets |
retry-after |
on a 429 only: seconds to wait. Never 0 |
x-request-id |
correlation id for this request; also in every error body |
The three rate-limit headers appear on every response where the credential produced a limit
verdict. x-ratelimit-reset-requests is a duration in seconds, not a timestamp.
Two spellings of the same window
The x-ratelimit-*-requests trio follows the OpenAI convention — note the -requests
suffix, which is easy to drop when writing a client from memory. The same numbers are also
published as the IETF draft pair RateLimit-Policy: "rpm";q=<limit>;w=60 and
RateLimit: "rpm";r=<remaining>;t=<seconds to reset>, whenever the limit came from the
shared cluster-wide window (and always on a 429). There is no RateLimit-Limit,
RateLimit-Remaining or RateLimit-Reset header.
Browser clients can read all of these headers cross-origin: they are named in the API's CORS
Access-Control-Expose-Headers, along with etag, location, link, x-resource-id and
preference-applied.
Two different 429s¶
A 429 means one of two things, and they are not interchangeable. Branch on code:
code |
Window | Waiting helps? |
|---|---|---|
rate_limited |
a rolling per-minute window | Yes — retry-after is at most 60 |
quota_exhausted |
your plan's daily allowance | Only until UTC midnight |
rate_limited — you are going too fast¶
{
"type": "https://platform.trilocore.com/docs/en/api/errors/#rate_limited",
"title": "Rate limit exceeded",
"status": 429,
"code": "rate_limited",
"error": "rate_limited",
"instance": "/api/v1/bevm/sessions/current/transactions",
"request_id": "req_…"
}
The response carries retry-after set to the exact remainder of the window — the real number of
seconds, not a flat 60 — plus x-ratelimit-remaining-requests: 0. Sleep for retry-after and
continue.
Your per-minute allowance comes from your plan and is reported in
x-ratelimit-limit-requests; do not hard-code it, because it changes when your plan does.
The limit still applies if the shared limiter is temporarily unavailable. The front door then
falls back to a stricter, conservative limit instead of none at all, so you may briefly see a
lower x-ratelimit-limit-requests and an earlier 429. The response looks the same: a
rate_limited 429 with retry-after and the three headers. Handle it the same way. This is
one more reason to read the limit from the headers instead of hard-coding it.
Contract search has its own per-minute budget¶
The contract-search endpoints — POST /api/v1/radar/code/search, search:count,
search:facets and GET /api/v1/radar/code/{code_id} — also draw on a second, separate budget
of 60 requests per minute per API key or signed-in user, on top of your plan's allowance. A
request must fit both. When the contract-search budget is the one that runs out, the 429 is the
same rate_limited document, and its retry-after and rate-limit headers describe that budget
(x-ratelimit-limit-requests: 60), so the same backoff code handles both.
Count several facets of one query with one search:facets request rather than one
search:count per facet: it answers every clause together and costs one request.
curl -s https://api.trilocore.ai/api/v1/radar/code/search:facets \
-H "Authorization: Bearer $TRILOCORE_API_KEY" -H 'content-type: application/json' \
-d '{"q": "opcodes: delegatecall", "facets": ["gates_present = false", "proxy_shape: none"]}'
Each entry of the answer is what search:count would return for (q) and (clause): a count,
whether it is exact (otherwise it is a lower bound), and the coverage of that count.
quota_exhausted — your daily allowance is spent¶
{
"type": "https://platform.trilocore.com/docs/en/api/errors/#quota_exhausted",
"title": "Daily quota exhausted",
"status": 429,
"code": "quota_exhausted",
"error": "quota_exhausted",
"used": 5000,
"limit": 5000,
"plan": "team",
"reset_at": "2026-08-16T00:00:00Z",
"instance": "/api/v1/bevm/sessions/current/transactions",
"request_id": "req_…"
}
Four extension members are merged into the problem document:
| Field | Meaning |
|---|---|
used |
requests consumed today |
limit |
the plan's daily allowance |
plan |
the plan the limit came from |
reset_at |
ISO-8601 instant when the allowance resets |
Quota resets at UTC midnight, and retry-after is set to the number of seconds until then.
That can be many hours, so this is not an error to sit and retry through: surface it, stop the
worker, or upgrade the plan. Retrying in a loop will consume nothing but your own CPU — a
rejected request is still rejected.
The unauthenticated endpoints¶
The four public verification endpoints are limited per client IP rather than per credential,
at 60 requests per minute, with separate buckets for the report and attestation surfaces. They
also cap request bodies more tightly than the authenticated API; an oversized body is rejected
with 413 payload_too_large rather than a 429.
Because these limits are keyed on IP, everyone behind a shared egress address shares a bucket. If you are verifying attestations in bulk from a fleet, spread the work or use an authenticated credential.
Handling limits well¶
Retry 429 on retry-after, and back off exponentially with jitter on 500 and 502. Nothing
else in the error catalogue should be retried unchanged.
import time
import httpx
def request_with_backoff(client: httpx.Client, method: str, url: str, **kw):
for attempt in range(5):
r = client.request(method, url, **kw)
if r.status_code == 429:
body = r.json()
if (body.get("code") or body.get("error")) == "quota_exhausted":
# Resets at UTC midnight — retrying is pointless. Fail loudly.
raise RuntimeError(f"daily quota spent, resets {body.get('reset_at')}")
time.sleep(int(r.headers.get("retry-after", "1")))
continue
if r.status_code in (500, 502):
time.sleep(2**attempt)
continue
return r
raise RuntimeError("giving up after 5 attempts")
const sleep = (s: number) => new Promise((r) => setTimeout(r, s * 1000));
async function requestWithBackoff(url: string, init: RequestInit) {
for (let attempt = 0; attempt < 5; attempt++) {
const r = await fetch(url, init);
if (r.status === 429) {
const body = await r.clone().json();
if ((body.code ?? body.error) === "quota_exhausted") {
// Resets at UTC midnight — retrying is pointless. Fail loudly.
throw new Error(`daily quota spent, resets ${body.reset_at}`);
}
await sleep(Number(r.headers.get("retry-after") ?? 1));
continue;
}
if (r.status === 500 || r.status === 502) {
await sleep(2 ** attempt);
continue;
}
return r;
}
throw new Error("giving up after 5 attempts");
}
Two habits keep you well clear of the limits:
- Poll conditionally. Send
If-None-Matchwith theETagyou were last given. A304still costs a request, but it is far cheaper for both sides than re-fetching an unchanged collection. - Reuse
Idempotency-Keyon retries. A retry that replays a stored response returns immediately, does no work, and cannot double-execute. See Idempotency.