Skip to content

Rate limits and quota

Every authenticated response tells you where you stand. Read the headers rather than counting requests yourself:

curl -sD - -o /dev/null https://api.trilocore.ai/api/v1/workspaces \
  -H "Authorization: Bearer $TRILOCORE_API_KEY" \
  | grep -i '^x-ratelimit\|^retry-after\|^x-request-id'
x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 597
x-ratelimit-reset-requests: 41
x-request-id: req_e345361146654a96898ff450881779d0

The headers

Header Meaning
x-ratelimit-limit-requests requests allowed in the current window
x-ratelimit-remaining-requests requests left in the current window
x-ratelimit-reset-requests seconds until the window resets
retry-after on a 429 only: seconds to wait. Never 0
x-request-id correlation id for this request; also in every error body

The three rate-limit headers appear on every response where the credential produced a limit verdict. x-ratelimit-reset-requests is a duration in seconds, not a timestamp.

Two spellings of the same window

The x-ratelimit-*-requests trio follows the OpenAI convention — note the -requests suffix, which is easy to drop when writing a client from memory. The same numbers are also published as the IETF draft pair RateLimit-Policy: "rpm";q=<limit>;w=60 and RateLimit: "rpm";r=<remaining>;t=<seconds to reset>, whenever the limit came from the shared cluster-wide window (and always on a 429). There is no RateLimit-Limit, RateLimit-Remaining or RateLimit-Reset header.

Browser clients can read all of these headers cross-origin: they are named in the API's CORS Access-Control-Expose-Headers, along with etag, location, link, x-resource-id and preference-applied.

Two different 429s

A 429 means one of two things, and they are not interchangeable. Branch on code:

code Window Waiting helps?
rate_limited a rolling per-minute window Yes — retry-after is at most 60
quota_exhausted your plan's daily allowance Only until UTC midnight

rate_limited — you are going too fast

{
  "type": "https://platform.trilocore.com/docs/en/api/errors/#rate_limited",
  "title": "Rate limit exceeded",
  "status": 429,
  "code": "rate_limited",
  "error": "rate_limited",
  "instance": "/api/v1/bevm/sessions/current/transactions",
  "request_id": "req_…"
}

The response carries retry-after set to the exact remainder of the window — the real number of seconds, not a flat 60 — plus x-ratelimit-remaining-requests: 0. Sleep for retry-after and continue.

Your per-minute allowance comes from your plan and is reported in x-ratelimit-limit-requests; do not hard-code it, because it changes when your plan does.

The limit still applies if the shared limiter is temporarily unavailable. The front door then falls back to a stricter, conservative limit instead of none at all, so you may briefly see a lower x-ratelimit-limit-requests and an earlier 429. The response looks the same: a rate_limited 429 with retry-after and the three headers. Handle it the same way. This is one more reason to read the limit from the headers instead of hard-coding it.

Contract search has its own per-minute budget

The contract-search endpoints — POST /api/v1/radar/code/search, search:count, search:facets and GET /api/v1/radar/code/{code_id} — also draw on a second, separate budget of 60 requests per minute per API key or signed-in user, on top of your plan's allowance. A request must fit both. When the contract-search budget is the one that runs out, the 429 is the same rate_limited document, and its retry-after and rate-limit headers describe that budget (x-ratelimit-limit-requests: 60), so the same backoff code handles both.

Count several facets of one query with one search:facets request rather than one search:count per facet: it answers every clause together and costs one request.

curl -s https://api.trilocore.ai/api/v1/radar/code/search:facets \
  -H "Authorization: Bearer $TRILOCORE_API_KEY" -H 'content-type: application/json' \
  -d '{"q": "opcodes: delegatecall", "facets": ["gates_present = false", "proxy_shape: none"]}'

Each entry of the answer is what search:count would return for (q) and (clause): a count, whether it is exact (otherwise it is a lower bound), and the coverage of that count.

quota_exhausted — your daily allowance is spent

{
  "type": "https://platform.trilocore.com/docs/en/api/errors/#quota_exhausted",
  "title": "Daily quota exhausted",
  "status": 429,
  "code": "quota_exhausted",
  "error": "quota_exhausted",
  "used": 5000,
  "limit": 5000,
  "plan": "team",
  "reset_at": "2026-08-16T00:00:00Z",
  "instance": "/api/v1/bevm/sessions/current/transactions",
  "request_id": "req_…"
}

Four extension members are merged into the problem document:

Field Meaning
used requests consumed today
limit the plan's daily allowance
plan the plan the limit came from
reset_at ISO-8601 instant when the allowance resets

Quota resets at UTC midnight, and retry-after is set to the number of seconds until then. That can be many hours, so this is not an error to sit and retry through: surface it, stop the worker, or upgrade the plan. Retrying in a loop will consume nothing but your own CPU — a rejected request is still rejected.

The unauthenticated endpoints

The four public verification endpoints are limited per client IP rather than per credential, at 60 requests per minute, with separate buckets for the report and attestation surfaces. They also cap request bodies more tightly than the authenticated API; an oversized body is rejected with 413 payload_too_large rather than a 429.

Because these limits are keyed on IP, everyone behind a shared egress address shares a bucket. If you are verifying attestations in bulk from a fleet, spread the work or use an authenticated credential.

Handling limits well

Retry 429 on retry-after, and back off exponentially with jitter on 500 and 502. Nothing else in the error catalogue should be retried unchanged.

import time

import httpx


def request_with_backoff(client: httpx.Client, method: str, url: str, **kw):
    for attempt in range(5):
        r = client.request(method, url, **kw)

        if r.status_code == 429:
            body = r.json()
            if (body.get("code") or body.get("error")) == "quota_exhausted":
                # Resets at UTC midnight — retrying is pointless. Fail loudly.
                raise RuntimeError(f"daily quota spent, resets {body.get('reset_at')}")
            time.sleep(int(r.headers.get("retry-after", "1")))
            continue

        if r.status_code in (500, 502):
            time.sleep(2**attempt)
            continue

        return r

    raise RuntimeError("giving up after 5 attempts")
const sleep = (s: number) => new Promise((r) => setTimeout(r, s * 1000));

async function requestWithBackoff(url: string, init: RequestInit) {
  for (let attempt = 0; attempt < 5; attempt++) {
    const r = await fetch(url, init);

    if (r.status === 429) {
      const body = await r.clone().json();
      if ((body.code ?? body.error) === "quota_exhausted") {
        // Resets at UTC midnight — retrying is pointless. Fail loudly.
        throw new Error(`daily quota spent, resets ${body.reset_at}`);
      }
      await sleep(Number(r.headers.get("retry-after") ?? 1));
      continue;
    }

    if (r.status === 500 || r.status === 502) {
      await sleep(2 ** attempt);
      continue;
    }

    return r;
  }
  throw new Error("giving up after 5 attempts");
}

Two habits keep you well clear of the limits:

  • Poll conditionally. Send If-None-Match with the ETag you were last given. A 304 still costs a request, but it is far cheaper for both sides than re-fetching an unchanged collection.
  • Reuse Idempotency-Key on retries. A retry that replays a stored response returns immediately, does no work, and cannot double-execute. See Idempotency.