Chat completions

Quickstart, authentication, streaming, routing and error handling for the Autonomous Relay Agents API. OpenAI-compatible: change the base URL and the key.

POST/v1/chat/completions

The main endpoint, and OpenAI-compatible. Request and response bodies match the OpenAI Chat Completions shape, so existing clients and SDKs work without a wrapper.

Parameters

  • model — a catalog id, author/slug. Required unless models is given.
  • messages — the conversation, as in OpenAI.
  • stream — server-sent events. See Streaming.
  • max_completion_tokens (or max_tokens) — an integer from 1 to 10,000,000; anything else is a 400 — temperature, top_p, top_k, stop, seed
  • frequency_penalty, presence_penalty, logit_bias, logprobs, top_logprobs; min_p and repetition_penalty where an open-weight host supports them
  • tools, tool_choice, parallel_tool_calls, response_format, reasoning
  • user — a stable id for your end user, passed to the provider for abuse attribution.
  • models — a fallback chain. See below.
  • provider — routing preferences. See Routing.

Message content is a string or a list of parts — text, image_url, input_audio, file, video_url. Tool calls and tool results work on every provider, not only the OpenAI-compatible ones. If an endpoint cannot accept one of your parts — audio sent to Anthropic, say — the request moves on to the next endpoint that can, and fails with a 400 if none can.

Unknown parameters are dropped, not rejected. Endpoints differ in what they support, and failing a request because one provider does not understand a field would make your code provider-specific — which is the thing you came here to avoid. A few OpenAI parameters are dropped on purpose, because they would change what you are billed or what the provider keeps — n, store, service_tier and web_search_options.

Model fallback

Pass models instead of model to give an ordered chain. If the first cannot be served, the next is tried, up to three attempts.

json
{
  "models": [
    "anthropic/claude-sonnet-5",
    "openai/gpt-4o",
    "google/gemini-2.0-flash"
  ],
  "messages": [{"role": "user", "content": "Hello"}]
}

The chain is walked to exhaustion rather than retried once, and every attempt is recorded so a failure is attributable afterwards.

Response

json
{
  "id": "gen-...",
  "model": "anthropic/claude-sonnet-5",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "Hello."},
    "finish_reason": "stop",
    "native_finish_reason": "end_turn"
  }],
  "usage": {"prompt_tokens": 9, "completion_tokens": 3, "total_tokens": 12}
}

native_finish_reason is the provider’s own value, preserved verbatim alongside the normalised one. Providers disagree about what “stop” means, and collapsing that distinction loses information you may need.

Usage is the provider’s own count, not our estimate. You are billed on what they reported.

App attribution

Send HTTP-Referer with your app’s URL and X-Title with its name, and each request is recorded against your app. X-App-Title and X-OpenRouter-Title are accepted as X-Title, so a client already set up for another router needs no change.

curl
curl -X POST https://autonomousrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $AGENTROUTER_API_KEY" \
  -H 'HTTP-Referer: https://myapp.example' \
  -H 'X-Title: My App' \
  -H 'content-type: application/json' \
  -d '{"model": "anthropic/claude-sonnet-5", "messages": [{"role": "user", "content": "Hello"}]}'
  • Only the scheme, host and path of the referer are kept. The query string and fragment are dropped before anything is stored.
  • X-App-Visibility: hidden marks the app private, so it will never appear in a public app listing. It is read on the app’s first request only.
  • X-App-Categories takes up to two comma-separated categories per request, and an app keeps up to ten.
  • A localhost or IP-address referer is recorded on the request but does not create an app.

Inspecting a request afterwards

GET/v1/generation?id=gen-...

Returns the metadata for one request: which endpoint served it, token counts, latency, cost, the price version applied and the app it came from. Never the content: this endpoint does not return it under any setting. If your organisation enables request logging — off by default, and not switchable yet — stored content is read from its own endpoint instead.