The main endpoint, and OpenAI-compatible. Request and response bodies match the OpenAI Chat Completions shape, so existing clients and SDKs work without a wrapper.
Parameters
model— a catalog id,author/slug. Required unlessmodelsis given.messages— the conversation, as in OpenAI.stream— server-sent events. See Streaming.max_completion_tokens(ormax_tokens) — an integer from 1 to 10,000,000; anything else is a 400 —temperature,top_p,top_k,stop,seedfrequency_penalty,presence_penalty,logit_bias,logprobs,top_logprobs;min_pandrepetition_penaltywhere an open-weight host supports themtools,tool_choice,parallel_tool_calls,response_format,reasoninguser— a stable id for your end user, passed to the provider for abuse attribution.models— a fallback chain. See below.provider— routing preferences. See Routing.
Message content is a string or a list of parts — text, image_url, input_audio, file, video_url. Tool calls and tool results work on every provider, not only the OpenAI-compatible ones. If an endpoint cannot accept one of your parts — audio sent to Anthropic, say — the request moves on to the next endpoint that can, and fails with a 400 if none can.
n, store, service_tier and web_search_options.Model fallback
Pass models instead of model to give an ordered chain. If the first cannot be served, the next is tried, up to three attempts.
{
"models": [
"anthropic/claude-sonnet-5",
"openai/gpt-4o",
"google/gemini-2.0-flash"
],
"messages": [{"role": "user", "content": "Hello"}]
}The chain is walked to exhaustion rather than retried once, and every attempt is recorded so a failure is attributable afterwards.
Response
{
"id": "gen-...",
"model": "anthropic/claude-sonnet-5",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "Hello."},
"finish_reason": "stop",
"native_finish_reason": "end_turn"
}],
"usage": {"prompt_tokens": 9, "completion_tokens": 3, "total_tokens": 12}
}native_finish_reason is the provider’s own value, preserved verbatim alongside the normalised one. Providers disagree about what “stop” means, and collapsing that distinction loses information you may need.
Usage is the provider’s own count, not our estimate. You are billed on what they reported.
App attribution
Send HTTP-Referer with your app’s URL and X-Title with its name, and each request is recorded against your app. X-App-Title and X-OpenRouter-Title are accepted as X-Title, so a client already set up for another router needs no change.
curl -X POST https://autonomousrelay.com/v1/chat/completions \
-H "Authorization: Bearer $AGENTROUTER_API_KEY" \
-H 'HTTP-Referer: https://myapp.example' \
-H 'X-Title: My App' \
-H 'content-type: application/json' \
-d '{"model": "anthropic/claude-sonnet-5", "messages": [{"role": "user", "content": "Hello"}]}'- Only the scheme, host and path of the referer are kept. The query string and fragment are dropped before anything is stored.
X-App-Visibility: hiddenmarks the app private, so it will never appear in a public app listing. It is read on the app’s first request only.X-App-Categoriestakes up to two comma-separated categories per request, and an app keeps up to ten.- A
localhostor IP-address referer is recorded on the request but does not create an app.
Inspecting a request afterwards
Returns the metadata for one request: which endpoint served it, token counts, latency, cost, the price version applied and the app it came from. Never the content: this endpoint does not return it under any setting. If your organisation enables request logging — off by default, and not switchable yet — stored content is read from its own endpoint instead.