Skip to main content

API Reference

The LLM Gateway provides two API formats: OpenAI-compatible and Anthropic-compatible.

Base URL​

https://llm.bankr.bot

Authentication​

All requests require a Bankr API key (bk_...) in the X-API-Key header or Authorization: Bearer token:

X-API-Key: bk_YOUR_API_KEY

or

Authorization: Bearer bk_YOUR_API_KEY

Generate API keys at bankr.bot/api-keys.


OpenAI-Compatible API​

List Models​

GET /v1/models

List available models, newest first. Deprecated and hidden models are not listed.

Response​

One entry shown — call /v1/models or bankr llm models for the full live catalog.

{
"object": "list",
"data": [
{
"id": "glm-5.3-flash",
"object": "model",
"name": "GLM-5.3 Flash",
"owned_by": "z-ai",
"context_window": 1048576,
"max_output_tokens": 131072,
"input_modalities": ["text", "image"],
"output_modalities": ["text"],
"private": true,
"attested": "gateway",
"pricing": {
"input": 0.15,
"output": 0.5,
"cache_read": 0.03,
"currency": "usd",
"unit": "million_tokens"
}
}
]
}

A model that can serve private (TEE) inference carries "private": true and "attested": "gateway"; one reachable with zero data retention carries "zdr": true. An image model reports "output_modalities": ["image"] and its image rate under pricing.image_output. pricing.discount appears only when a discount applies to you.


Chat Completions​

POST /v1/chat/completions

Create a chat completion using OpenAI format.

Request​

curl -X POST https://llm.bankr.bot/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"temperature": 0.7,
"max_tokens": 1024
}'

Response​

{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1706123456,
"model": "claude-opus-4.8",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 10,
"total_tokens": 30
}
}

Image Generation​

POST /v1/images/generations

Generate images using OpenAI's native Images API format (today: the gpt-image-2.5 line and gpt-image-2). Streaming is not supported and n is capped at 4.

Request​

curl -X POST https://llm.bankr.bot/v1/images/generations \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "gpt-image-2.5-flare",
"prompt": "a friendly pixel-art robot mascot holding a coin",
"size": "1024x1024",
"n": 1
}'

Response​

Images are returned as base64 in b64_json; usage reports the tokens billed.

{
"created": 1783526478,
"data": [{ "b64_json": "iVBORw0KGgoAAAANSU..." }],
"usage": {
"input_tokens": 19,
"output_tokens": 1756,
"output_tokens_details": { "image_tokens": 1756, "text_tokens": 0 },
"total_tokens": 1775
}
}

See the Image Generation guide for SDK examples, parameters, and pricing.


Anthropic-Compatible API​

Messages​

POST /v1/messages

Create a message using Anthropic format. Ideal for Claude Code and Anthropic SDK users.

Request​

curl -X POST https://llm.bankr.bot/v1/messages \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello!"}
]
}'

Response​

{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
],
"model": "claude-opus-4.8",
"stop_reason": "end_turn",
"usage": {
"input_tokens": 10,
"output_tokens": 12
}
}

Ignored Request Fields​

When a request is served through OpenRouter's native /messages, the OpenRouter routing controls — provider, models, fallbacks, route, transforms, plugins — are removed rather than forwarded. The gateway picks the upstream provider itself, and usage is metered and priced against the model you named, so a routing override in the body can't move the request to a model you aren't being billed for.

Don't send them at all, though: that route is the only one that strips them, and a provider that receives an unknown top-level field may reject the whole request.


Privacy Tiers​

Every request is served at standard (default), zdr or private, chosen with the "privacy" body field, a /zdr or /private base-path prefix, a :zdr or :private model suffix, or the account setting. Tiers fail closed rather than downgrading. See Privacy Tiers for the rules, error codes and the X-Privacy-Tier response header, and Private Inference for the enclave's attestation headers.

Attestation Report​

GET /v1/attestation/report

Returns the provider's raw attestation report (Intel TDX quote + ed25519 signing key) so a client can independently verify the enclave. See Private Inference for the full flow.


Health Check​

GET /health

Check gateway and provider health. No authentication required.

Response​

{
"status": "ok",
"providers": {
"vertexGemini": true,
"vertexClaude": true,
"openrouter": true
}
}

Status codes:

  • 200 — At least one provider healthy
  • 503 — All providers unavailable

Error Responses​

{
"error": {
"message": "Model temporarily unavailable",
"type": "api_error",
"code": "provider_unavailable"
}
}
Statustype / codeMeaning
400invalid_request_errorInvalid request — e.g. an unknown model (unsupported_model) or a missing field (missing_required_fields)
401auth_errorAPI key missing, invalid or inactive
402insufficient_credits or daily_budget_exceededCredits exhausted, or the daily spend budget is reached
403auth_errorLLM Gateway isn't enabled on the key
410model_deprecated / model_removedThe model is past its retirement date — the message names the replacement
422invalid_request_errorThe model can't be served at the requested privacy tier
429rate_limit_errorMore than 60 requests a minute from one API key or one IP
503api_error / provider_unavailableNo provider is serving the model right now — retry, or fall back to another model
504timeout_error / request_timeoutThe upstream provider timed out
500api_error / internal_errorUnexpected gateway error

An error returned by the upstream provider keeps its status and message. On /v1/messages, errors use the Anthropic envelope instead — {"type": "error", "error": {"type", "message"}} — with the same status codes.


Usage​

Get Usage Summary​

GET /v1/usage?days=30

Returns aggregated token usage and cost breakdown for the authenticated API key. Requires authentication.

Query Parameters​

ParameterTypeDefaultDescription
daysnumber30Number of days to aggregate (1–90)

Response​

{
"object": "usage_summary",
"days": 30,
"startDate": "2026-01-28T00:00:00.000Z",
"endDate": "2026-02-27T00:00:00.000Z",
"totals": {
"totalRequests": 1981,
"totalInputTokens": 489789,
"totalOutputTokens": 631794,
"totalCacheReadInputTokens": 53460194,
"totalCacheWriteInputTokens": 12555591,
"totalTokens": 67137368,
"totalCost": 248.38,
"totalCacheCost": 208.01
},
"byModel": [
{
"model": "claude-opus-4.8",
"provider": "vertex-claude",
"requests": 1574,
"inputTokens": 6250,
"outputTokens": 500097,
"cacheReadInputTokens": 27309491,
"cacheWriteInputTokens": 7474005,
"totalTokens": 35289843,
"totalCost": 218.7,
"cacheCost": 181.1
}
]
}

Credits​

Get Credit Balance​

GET /v1/credits

Returns the current LLM credit balance for the API key's wallet. Requires authentication.

Use this to check available capacity before relying on the gateway — effectiveBalanceUsd is the truest "available balance" because it nets out in-flight usage that hasn't been deducted yet. Balances read directly from the database (not a cached value) for accuracy.

Request​

curl https://llm.bankr.bot/v1/credits \
-H "X-API-Key: bk_YOUR_API_KEY"

Response​

{
"object": "credit_balance",
"balanceUsd": 12.34,
"effectiveBalanceUsd": 11.2,
"undeductedCostUsd": 1.14,
"dailyBudget": {
"limitUsd": 25,
"spentUsd": 4.2,
"remainingUsd": 20.8,
"exceeded": false,
"windowHours": 24
}
}

Fields​

FieldTypeDescription
balanceUsdnumberTotal spendable credit on the wallet, in USD.
effectiveBalanceUsdnumberAvailable balance after subtracting in-flight usage not yet deducted. Floored at 0. Use this for capacity decisions.
undeductedCostUsdnumberCost of in-flight/served requests not yet deducted from balanceUsd (the amount subtracted to derive effectiveBalanceUsd).
dailyBudgetobjectPresent only when a daily spend budget is set on the wallet. Omitted entirely when spend is uncapped.

Requests are rejected with 402 Payment Required once the balance runs out. A chat or messages request is also rejected when its worst-case cost (the prompt plus max_tokens or max_completion_tokens, or the model's maximum output if you set neither) is more than the balance left after your other in-flight requests; a lower max_tokens lets it through. There are no per-key spending caps — the balance shown is the full credit available to the key.

Daily spend budget​

A wallet can carry an optional cap on how much it spends in any rolling 24-hour window. It bounds burn rate; the credit balance still bounds total spend. Set it in the Settings tab at bankr.bot/terminal/llm — a key cannot change its own wallet's budget, so a leaked key can't raise the cap it is bound by.

The window trails the current moment rather than resetting on a clock boundary, so the cap holds over every 24 hours rather than per calendar day — there is no midnight at which a spent budget returns in full. Capacity comes back gradually, as individual charges age past 24 hours old.

The budget spans all metered LLM spend on the wallet, not just gateway traffic: requests from every API key it owns, plus Max Mode and app-invoked Bankr agent runs. Both are enforced — gateway requests are refused, and an agent run ends early once spend in the window reaches the cap.

FieldTypeDescription
limitUsdnumberThe configured cap, in USD per rolling 24-hour window.
spentUsdnumberSpend counted against the budget in the trailing 24 hours, including usage not yet deducted.
remainingUsdnumberlimitUsd - spentUsd, floored at 0.
exceededbooleanWhether the budget is spent. While true, requests are rejected.
windowHoursnumberHours the window spans. Always 24.

Once the budget is reached, requests are rejected with 402 Payment Required and an error type of daily_budget_exceeded — distinct from the insufficient_credits returned when the balance itself runs out. Only requests that spend are blocked. Every read-only (GET) endpoint — /v1/credits, /v1/usage, /v1/models among them — keeps working while you're over budget, so poll exceeded on /v1/credits to know when you're unblocked — there is no reset time to schedule against, since capacity returns as individual charges age out — and check /v1/usage for what spent the budget:

{
"error": {
"message": "Daily LLM Gateway spend budget reached. It covers a rolling 24 hours, so it frees up as earlier usage ages out, or you can raise it at bankr.bot/llm.",
"type": "daily_budget_exceeded"
}
}

Budget changes reach the gateway within 60 seconds (it caches authentication state), and enforcement is evaluated against that cached view, so spend may overshoot the cap slightly under a sustained burst. Treat it as a guardrail, not an accounting boundary — your credit balance remains the hard limit on total spend.


Streaming​

Both endpoints support streaming responses:

curl -X POST https://llm.bankr.bot/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'

Streaming uses Server-Sent Events (SSE) format.