Limits & credits

Quickstart, authentication, streaming, routing and error handling for the Autonomous Relay Agents API. OpenAI-compatible: change the base URL and the key.

Rate limits

Request rate is keyed on lifetime purchase rather than current balance, so topping up raises your ceiling and spending down does not lower it mid-workload. Limits are shared across instances, so they are the same number whether one process is calling or twenty.

Your provider’s own rate limits also apply, and are frequently the binding constraint — you are their customer.

Spend ceilings

Set a limit per key and per workspace. A key’s limit covers a window — daily, weekly, monthly or lifetime — and has a mode. A hard limit, the default, refuses a request once the key is at or past it, with 403 budget_exceeded before any provider is asked. The refusal is recorded, so it appears in Activity with the reason: you can see what was stopped, not just what ran. A soft limit serves the request and adds an x-agentsrouter-limit-warning header carrying the same sentence.

GET /v1/keys returns each capped key’s usage in its current window, limit_remaining and limit_resets_at, measured exactly as the limit is enforced. A limit can be changed on a live key without re-issuing it.

PATCH/v1/keys/:hash
curl
curl -X PATCH https://autonomousrelay.com/v1/keys/$KEY_HASH \
  -H "Authorization: Bearer $AGENTROUTER_MANAGEMENT_KEY" \
  -H 'content-type: application/json' \
  -d '{"limit": 50, "limit_reset": "monthly", "limit_mode": "hard"}'
GET/v1/workspaces/:id/budgets
PUT/v1/workspaces/:id/budgets/:interval

Clients (workspaces)

Keep one workspace per client. Each has its own keys, its own budgets and its own spend, under your one account and balance. A workspace budget bounds that workspace’s spend alone, so one client running hot never stops another. Create a key in a workspace with workspace_id on POST /v1/keys, or move one with PATCH /v1/keys/:hash.

GET/v1/workspaces
POST/v1/workspaces
PATCH/v1/workspaces/:id
DELETE/v1/workspaces/:id

To re-bill a client, export its month as CSV: one row per request, with billed_usd (what you were charged) and list_price_usd (the tokens at list price, which is what budgets measure). /v1/activity and /v1/activity/summary take the same workspace filter.

GET/v1/activity/export.csv?month=YYYY-MM&workspace=:id

Client reports

Each client (workspace) has a printable monthly report — requests, tokens and usage value by model, agent and day — under your own name, logo and footer rather than ours. Set the branding once; the logo is a PNG or JPEG under 150 KB, checked by its bytes and never fetched from a URL. Usage value is the catalogue list price of the tokens used: a usage report, not an invoice.

GET/v1/workspaces/:id/report?month=YYYY-MM
PUT/v1/branding

Members and roles

Several people can share an organisation, each with a role. Owner: everything, including inviting and removing people. Developer: keys, provider keys, limits, clients, alerts and agents; reads billing. Finance: activity, exports, credits and buying credits; changes nothing else. A request a role does not allow is refused with 403 permission_denied. A management key with no person behind it has full power. Invitations are single-use links that expire in seven days; a person can belong to several organisations and switch between them.

GET/v1/members
POST/v1/invitations
POST/v1/auth/switch

Agents

Name the agent on each request with X-AgentsRouter-Agent (1 to 64 letters, digits, . _ : -). Its spend is tracked across every key it uses, and a budget on it refuses its requests with 403 budget_exceeded once reached — without touching other agents on the same key. A soft budget serves instead and adds an x-agentsrouter-limit-warning header. Either way, crossing 50%, 80% and 100% of a budget is announced on your alert channels as agent_limit.threshold. A malformed name is a 400 before anything is spent. A budget may also be scoped to one client workspace with workspace_id: it then binds only requests from that workspace’s keys, and an organisation-wide budget on the same agent still applies alongside it.

GET/v1/agents
PUT/v1/agents/:agent/budgets/:interval
DELETE/v1/agents/:agent/budgets/:interval

Usage alerts

Add a Slack, webhook or email channel, and each key with a spend limit is announced when it crosses 50%, 80% and 100% of it — or the thresholds you choose — once per threshold per limit window. A key that jumps several thresholds in one request is announced once, at the highest. Every attempt is logged with its outcome.

POST/v1/alerts/channels
POST/v1/alerts/channels/:id/test
GET/v1/alerts/deliveries

Runaway spend is flagged without any limit: every five minutes, a key or an agent whose last 15 minutes cost at least 3× its usual 15 minutes (and at least $5), or made at least 5× its usual requests (and at least 300), is flagged once for that hour and sent to every channel as a spend.runaway event. “Usual” is its own previous 24 hours. By default nothing is paused: you decide what to do. Set runaway_action to pause and a flagged key is also paused: it answers 403 key_paused until you resume it with PATCH /v1/keys/:hash {"paused": false} or from the Alerts page. An agent is never paused: a tag is not a credential.

GET/v1/alerts/runaways
GET/v1/alerts/settings
PUT/v1/alerts/settings

A webhook delivery is JSON with an x-agentsrouter-signature: t=<unix>,v1=<hex> header. Verify it with the signing_secret returned once when the channel is created: v1 is HMAC-SHA256 of <t>.<raw body>. Targets must be public https URLs; a Slack target must be a hooks.slack.com incoming webhook.

Credits

GET/v1/credits
GET/v1/credits/history

Balance, plus what is held in flight. A hold is placed before dispatch and the unused remainder is returned at settlement, so concurrent requests cannot collectively overspend a balance that looked sufficient to each of them individually.

A request is refused when it would take the balance below zero — or below credit_line, if an Enterprise agreement has set one. That is a check on admission, not a guarantee about the balance: settlement uses the provider’s reported usage, which can exceed what was reserved, and a stream already in flight cannot be cut off part-way. A balance can therefore end up negative, and while it is, every request is refused — free models included — until it is topped up.

Usage

GET/v1/activity

Per-request metadata: model, provider, tokens, latency, cost, outcome. Paginated with a cursor. Never the content of your requests: this endpoint does not return it under any setting, and it is not stored at all unless your organisation enables request logging — off by default, and not switchable yet.

Each row also carries its routing overhead: the request’s total time minus the time spent inside providers, which is the latency this gateway itself added.

GET/v1/activity/summary

Totals for the window, with the period of equal length before it alongside: spend, requests, token volume, cache hit rate, and blended dollars per million. Token volume is the same total_tokens each response reported — cached and reasoning tokens are counted once, inside the prompt and completion figures.

Cache hit rate and blended cost are null, not zero, when there is nothing to divide by. A rate of zero and no traffic at all are different answers.

GET/v1/activity/upstream

One row per dispatch to a provider. A request that failed over appears once per attempt, so a retry reads as the two upstream requests it actually was. Pass a request id to see one generation’s attempts in cascade order.

GET/v1/activity/sessions

Generations grouped by session_id, most recently active first, with every model the conversation was served by. More than one means it switched model mid-session, which is what breaks a warm prompt cache.