API Reference
The LLM Gateway provides two API formats: OpenAI-compatible and Anthropic-compatible.
Base URL
https://llm.bankr.bot
Authentication
All requests require a Bankr API key (bk_...) in the X-API-Key header or Authorization: Bearer token:
X-API-Key: bk_YOUR_API_KEY
or
Authorization: Bearer bk_YOUR_API_KEY
Generate API keys at bankr.bot/api-keys.
OpenAI-Compatible API
List Models
GET /v1/models
List available models, newest first. Deprecated and hidden models are not listed.
Response
One entry shown — call /v1/models or bankr llm models for the full live catalog.
{
"object": "list",
"data": [
{
"id": "glm-5.3-flash",
"object": "model",
"name": "GLM-5.3 Flash",
"owned_by": "z-ai",
"context_window": 1048576,
"max_output_tokens": 131072,
"input_modalities": ["text", "image"],
"output_modalities": ["text"],
"private": true,
"attested": "gateway",
"pricing": {
"input": 0.15,
"output": 0.5,
"cache_read": 0.03,
"currency": "usd",
"unit": "million_tokens"
}
}
]
}
A model that can serve private (TEE) inference carries "private": true and "attested": "gateway"; one reachable with zero data retention carries "zdr": true. An image model reports "output_modalities": ["image"] and its image rate under pricing.image_output. pricing.discount appears only when a discount applies to you.
Chat Completions
POST /v1/chat/completions
Create a chat completion using OpenAI format.
Request
curl -X POST https://llm.bankr.bot/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"temperature": 0.7,
"max_tokens": 1024
}'
Response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1706123456,
"model": "claude-opus-4.8",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 10,
"total_tokens": 30
}
}
Image Generation
POST /v1/images/generations
Generate images using OpenAI's native Images API format (today: the gpt-image-2.5 line and gpt-image-2). Streaming is not supported and n is capped at 4.
Request
curl -X POST https://llm.bankr.bot/v1/images/generations \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "gpt-image-2.5-flare",
"prompt": "a friendly pixel-art robot mascot holding a coin",
"size": "1024x1024",
"n": 1
}'
Response
Images are returned as base64 in b64_json; usage reports the tokens billed.
{
"created": 1783526478,
"data": [{ "b64_json": "iVBORw0KGgoAAAANSU..." }],
"usage": {
"input_tokens": 19,
"output_tokens": 1756,
"output_tokens_details": { "image_tokens": 1756, "text_tokens": 0 },
"total_tokens": 1775
}
}
See the Image Generation guide for SDK examples, parameters, and pricing.
Anthropic-Compatible API
Messages
POST /v1/messages
Create a message using Anthropic format. Ideal for Claude Code and Anthropic SDK users.
Request
curl -X POST https://llm.bankr.bot/v1/messages \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
Response
{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
],
"model": "claude-opus-4.8",
"stop_reason": "end_turn",
"usage": {
"input_tokens": 10,
"output_tokens": 12
}
}
Ignored Request Fields
When a request is served through OpenRouter's native /messages, the OpenRouter routing controls — provider, models, fallbacks, route, transforms, plugins — are removed rather than forwarded. The gateway picks the upstream provider itself, and usage is metered and priced against the model you named, so a routing override in the body can't move the request to a model you aren't being billed for.
Don't send them at all, though: that route is the only one that strips them, and a provider that receives an unknown top-level field may reject the whole request.
Privacy Tiers
Every request is served at standard (default), zdr or private, chosen with the "privacy" body field, a /zdr or /private base-path prefix, a :zdr or :private model suffix, or the account setting. Tiers fail closed rather than downgrading. See Privacy Tiers for the rules, error codes and the X-Privacy-Tier response header, and Private Inference for the enclave's attestation headers.
Attestation Report
GET /v1/attestation/report
Returns the provider's raw attestation report (Intel TDX quote + ed25519 signing key) so a client can independently verify the enclave. See Private Inference for the full flow.
Health Check
GET /health
Check gateway and provider health. No authentication required.
Response
{
"status": "ok",
"providers": {
"vertexGemini": true,
"vertexClaude": true,
"openrouter": true
}
}
Status codes:
200— At least one provider healthy503— All providers unavailable
Error Responses
{
"error": {
"message": "Model temporarily unavailable",
"type": "api_error",
"code": "provider_unavailable"
}
}
| Status | type / code | Meaning |
|---|---|---|
400 | invalid_request_error | Invalid request — e.g. an unknown model (unsupported_model) or a missing field (missing_required_fields) |
401 | auth_error | API key missing, invalid or inactive |
402 | insufficient_credits or daily_budget_exceeded | Credits exhausted, or the daily spend budget is reached |
403 | auth_error | LLM Gateway isn't enabled on the key |
410 | model_deprecated / model_removed | The model is past its retirement date — the message names the replacement |
422 | invalid_request_error | The model can't be served at the requested privacy tier |
429 | rate_limit_error | More than 60 requests a minute from one API key or one IP |
503 | api_error / provider_unavailable | No provider is serving the model right now — retry, or fall back to another model |
504 | timeout_error / request_timeout | The upstream provider timed out |
500 | api_error / internal_error | Unexpected gateway error |
An error returned by the upstream provider keeps its status and message. On /v1/messages, errors use the Anthropic envelope instead — {"type": "error", "error": {"type", "message"}} — with the same status codes.
Usage
Get Usage Summary
GET /v1/usage?days=30
Returns aggregated token usage and cost breakdown for the authenticated API key. Requires authentication.
Query Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
days | number | 30 | Number of days to aggregate (1–90) |
Response
{
"object": "usage_summary",
"days": 30,
"startDate": "2026-01-28T00:00:00.000Z",
"endDate": "2026-02-27T00:00:00.000Z",
"totals": {
"totalRequests": 1981,
"totalInputTokens": 489789,
"totalOutputTokens": 631794,
"totalCacheReadInputTokens": 53460194,
"totalCacheWriteInputTokens": 12555591,
"totalTokens": 67137368,
"totalCost": 248.38,
"totalCacheCost": 208.01
},
"byModel": [
{
"model": "claude-opus-4.8",
"provider": "vertex-claude",
"requests": 1574,
"inputTokens": 6250,
"outputTokens": 500097,
"cacheReadInputTokens": 27309491,
"cacheWriteInputTokens": 7474005,
"totalTokens": 35289843,
"totalCost": 218.7,
"cacheCost": 181.1
}
]
}
Credits
Get Credit Balance
GET /v1/credits
Returns the current LLM credit balance for the API key's wallet. Requires authentication.
Use this to check available capacity before relying on the gateway — effectiveBalanceUsd is the truest "available balance" because it nets out in-flight usage that hasn't been deducted yet. Balances read directly from the database (not a cached value) for accuracy.
Request
curl https://llm.bankr.bot/v1/credits \
-H "X-API-Key: bk_YOUR_API_KEY"
Response
{
"object": "credit_balance",
"balanceUsd": 12.34,
"effectiveBalanceUsd": 11.2,
"undeductedCostUsd": 1.14,
"dailyBudget": {
"limitUsd": 25,
"spentUsd": 4.2,
"remainingUsd": 20.8,
"exceeded": false,
"windowHours": 24
}
}
Fields
| Field | Type | Description |
|---|---|---|
balanceUsd | number | Total spendable credit on the wallet, in USD. |
effectiveBalanceUsd | number | Available balance after subtracting in-flight usage not yet deducted. Floored at 0. Use this for capacity decisions. |
undeductedCostUsd | number | Cost of in-flight/served requests not yet deducted from balanceUsd (the amount subtracted to derive effectiveBalanceUsd). |
dailyBudget | object | Present only when a daily spend budget is set on the wallet. Omitted entirely when spend is uncapped. |
Requests are rejected with 402 Payment Required once the balance runs out. A chat or messages request is also rejected when its worst-case cost (the prompt plus max_tokens or max_completion_tokens, or the model's maximum output if you set neither) is more than the balance left after your other in-flight requests; a lower max_tokens lets it through. There are no per-key spending caps — the balance shown is the full credit available to the key.
Daily spend budget
A wallet can carry an optional cap on how much it spends in any rolling 24-hour window. It bounds burn rate; the credit balance still bounds total spend. Set it in the Settings tab at bankr.bot/terminal/llm — a key cannot change its own wallet's budget, so a leaked key can't raise the cap it is bound by.
The window trails the current moment rather than resetting on a clock boundary, so the cap holds over every 24 hours rather than per calendar day — there is no midnight at which a spent budget returns in full. Capacity comes back gradually, as individual charges age past 24 hours old.
The budget spans all metered LLM spend on the wallet, not just gateway traffic: requests from every API key it owns, plus Max Mode and app-invoked Bankr agent runs. Both are enforced — gateway requests are refused, and an agent run ends early once spend in the window reaches the cap.
| Field | Type | Description |
|---|---|---|
limitUsd | number | The configured cap, in USD per rolling 24-hour window. |
spentUsd | number | Spend counted against the budget in the trailing 24 hours, including usage not yet deducted. |
remainingUsd | number | limitUsd - spentUsd, floored at 0. |
exceeded | boolean | Whether the budget is spent. While true, requests are rejected. |
windowHours | number | Hours the window spans. Always 24. |
Once the budget is reached, requests are rejected with 402 Payment Required and an error type of daily_budget_exceeded — distinct from the insufficient_credits returned when the balance itself runs out. Only requests that spend are blocked. Every read-only (GET) endpoint — /v1/credits, /v1/usage, /v1/models among them — keeps working while you're over budget, so poll exceeded on /v1/credits to know when you're unblocked — there is no reset time to schedule against, since capacity returns as individual charges age out — and check /v1/usage for what spent the budget:
{
"error": {
"message": "Daily LLM Gateway spend budget reached. It covers a rolling 24 hours, so it frees up as earlier usage ages out, or you can raise it at bankr.bot/llm.",
"type": "daily_budget_exceeded"
}
}
Budget changes reach the gateway within 60 seconds (it caches authentication state), and enforcement is evaluated against that cached view, so spend may overshoot the cap slightly under a sustained burst. Treat it as a guardrail, not an accounting boundary — your credit balance remains the hard limit on total spend.
Streaming
Both endpoints support streaming responses:
curl -X POST https://llm.bankr.bot/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-API-Key: bk_YOUR_API_KEY" \
-d '{
"model": "claude-opus-4.8",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'
Streaming uses Server-Sent Events (SSE) format.