Changelog
Dated changes to the Autonomous Relay Agents API and model catalogue. New endpoints, routing changes and catalogue updates that affect how you integrate.
- Added — Classifiers (beta): tag requests, and route them through a preset: Rules that tag each request from what it is — the model it names, its key, client, agent and app, whether it carries tools or images, streaming, and its size — never from what the prompt says. Tags are returned in x-agentsrouter-tags, stored on the request and filterable with GET /v1/activity?tag=. A rule can also send matching requests through a preset; a request naming its own @preset keeps it. /v1/classifiers, with an evaluate action to try a sample request.
- Added — Server tools: datetime and model search, run by the gateway: Declare {"type": "agentrouter:datetime"} or {"type": "agentrouter:experimental__search_models"} in tools[]: when the model calls one, the gateway runs it and feeds back the result, and your application receives only the final answer. Both are free; usage is reported in usage.server_tool_use and capped by max_tool_calls. Non-streaming requests only for now. GET and PUT /v1/tools list the catalogue and turn tools off per organisation; a tool that is off is refused with a 403.
- Added — Presets: versioned request settings, used as @preset/name: A preset is a named, versioned set of request settings — model or fallback models, temperature, top_p, top_k, max_tokens, seed, provider preferences, reasoning — that a request uses by sending "model": "@preset/<name>" (or @preset/<name>@<version> to pin one). The request’s own fields win. Every change is a new version and any version can be made current again. Settings only: messages and system prompts are not stored. /v1/presets and a Presets page in the dashboard.
- Added — Broadcast: each request’s record to your webhook or OpenTelemetry endpoint: POST /v1/broadcast adds a destination that receives a record of every request — model, provider, outcome, tokens, cost, timing, key name, client, agent, session and run, never the prompt or completion. A webhook gets signed JSON; an OpenTelemetry (OTLP/HTTP) endpoint such as Honeycomb or Langfuse gets one span per request with gen_ai attributes, and its API-key header is stored encrypted. Sent after the response, with sampling and a test action; the dashboard has an Observability page.
- Added — Routing defaults for the account: a default model, fallbacks and the provider sort: PUT /v1/routing sets, for every key in the organisation, a default model used when a request names none (instead of a 400), fallback models tried after the ones a request names, and the provider sort when a request sets none — lockable so a request cannot change it. Every model still passes guardrails, the privacy policy and BYOK. Owners and Admins only; the dashboard has a Routing page.
- Changed — A guardrail’s enforce_zdr is now enforced: A guardrail’s per-group enforce_zdr was stored and returned but not applied when routing. It is now: a key under such a guardrail is served only by endpoints that declare zero data retention for that model group, and is refused with a 404 naming the requirement otherwise. If you set enforce_zdr, check the models those keys use.
- Added — An account privacy policy, with an eligibility preview: PUT /v1/privacy sets requirements for every key in the organisation: zero data retention for every model or per model group, only endpoints that declare they do not train on prompts, and data regions. Each can only narrow routing, on top of guardrails, and a request cannot relax it. POST /v1/privacy/preview counts the models a policy leaves before you save it. Owners and Admins only; the dashboard has a Privacy page.
- Added — GET and PATCH /v1/org: the organisation, and renaming it: GET /v1/org returns the organisation’s id, name, slug, plan, creation date and number of seats; any role may read it. PATCH /v1/org {"name"} renames it, for Owners and Admins only. The dashboard has a Settings page for both.
- Changed — Runs can be read and stopped with a management key: GET /v1/runs, GET /v1/runs/{id}, and the close and resume actions now also accept a management key or a dashboard session, held to the person’s role: usage:read to look, config:write to close or resume. Opening a run still takes an inference key. GET /v1/runs takes ?status=, and a paused run can now be closed as well as resumed. The dashboard has a Runs page.
- Added — GET /v1/activity/{id}: one request in full: With a management key: the request’s routing and timing, the key that made it (by name and prefix), its client, agent, session and run, the usage lines it was billed on, and every provider attempt in order. The dashboard’s new Logs page opens it from any row.
- Changed — The API moved to autonomousrelay.com: The base URL is now https://autonomousrelay.com/v1, or https://api.autonomousrelay.com/v1. Requests to the old isync.ai addresses are redirected, but HTTP clients drop the Authorization header on a redirect to another host, so update your base URL. Keys, models and everything else are unchanged.
- Changed — Cached tokens are valued at the provider’s cache rate: Anthropic and OpenAI models now carry their published cache-read prices, and Anthropic models their 5-minute and 1-hour cache-write prices. A cached prompt token is counted at that rate in spend ceilings, budgets and activity, instead of the full input price. Each endpoint’s prices are listed by GET /v1/models/{author}/{slug}/endpoints.
- Added — Agent budgets per client workspace: PUT /v1/agents/{agent}/budgets/{interval} accepts workspace_id. The budget then binds only requests from that workspace’s keys, measured as that workspace’s spend for the agent, and an organisation-wide budget on the same agent still applies beside it. DELETE takes the same workspace_id as a query parameter.
- Added — A runaway-spend flag can pause the key: Set runaway_action to pause with PUT /v1/alerts/settings, and a key flagged for runaway spend is paused: it answers 403 key_paused until it is resumed with PATCH /v1/keys/{hash} and {"paused": false}, or from the Alerts or Keys page. A key can also be paused by hand the same way. Off by default, and an agent flag never pauses anything.
- Changed — Anthropic models take reasoning as extended thinking: On Anthropic models, reasoning with effort or max_tokens now turns on extended thinking with that budget, and a non-streaming reply returns the thinking as message.reasoning, as streaming already did with delta.reasoning. temperature, top_p and top_k are dropped while thinking is on, because Anthropic refuses them together.
- Added — Spend and budgets per agent: Name the agent on each request with the X-AgentsRouter-Agent header, and its spend is tracked across every key it uses. GET /v1/agents lists agents with their spend this month, and PUT /v1/agents/{agent}/budgets/{interval} sets a hard budget, which refuses with 403 budget_exceeded, or a soft one, which serves and adds a warning header.
- Added — Usage alerts and runaway-spend flags: Add a webhook, Slack or email channel with POST /v1/alerts/channels and be told when a key or an agent crosses 50, 80 or 100 percent of its spend limit. A key or agent whose last 15 minutes are far above its own previous 24 hours is flagged as spend.runaway on the same channels, and GET /v1/alerts/runaways lists the flags.
- Added — Client workspaces, exports and reports: Group keys into one workspace per client, each with its own budget, through /v1/workspaces. GET /v1/activity/export.csv exports a month of activity for one client, and GET /v1/workspaces/{id}/report returns a monthly client report that carries your own name, logo and footer.
- Added — Team seats and roles: Invite people with POST /v1/invitations and give each one a role: Owner, Admin, Developer, Finance or Read-only. Every management route checks the role of the person behind the dashboard session, so Finance can read billing without changing configuration.
- Changed — Failed requests say who failed: The activity log now records a router.* code in error_code when this gateway is what failed — a credential it could not open, a request it could not express — and leaves error_code to the provider’s own code when the provider answered. An attempt that broke mid-answer carries its reason on the upstream row. The errors count in the activity summary counts requests that ended with a code; before this it counted none, because no code was recorded.
- Changed — A model’s deprecation date is a calendar day: deprecation_date is sent as 2026-09-30 rather than as a timestamp at midnight UTC. A date is a day, and read as an instant it named the day before for anyone west of Greenwich.
- Added — Provider policy links in the catalogue: The provider directory now publishes each upstream’s privacy policy, terms of service and status page, where we could confirm them. Where we could not, the field stays empty rather than carrying a guess.
- Changed — The API description lists both production origins: openapi.json names https://isync.ai and https://api.isync.ai as its servers, so a client generated from the document targets the host it was fetched from.
- Changed — api.isync.ai serves the API only: The API hostname no longer returns the marketing site. Paths outside the API redirect to the documentation, so a mistyped path is visible as a mistake instead of answering with a page.
- Added — Preview a route before sending it: POST /v1/route/preview shows which endpoint would serve a request, and why each other candidate was set aside, without sending it or spending anything.