Configuration¶
Environment variables¶
Required¶
| Variable | Description |
|---|---|
PROXY_TOKEN |
UUID token generated by the platform in Settings → LLM Mode → Proxy. Validated on every request with constant-time comparison. |
Multi-token auth (hot rotation)¶
To accept multiple tokens simultaneously — useful when rotating keys without downtime — add numbered variants. Any listed token is accepted:
PROXY_TOKEN=token-primary
PROXY_TOKEN_1=token-secondary # add new token
PROXY_TOKEN_2=token-backup
# Remove PROXY_TOKEN later; PROXY_TOKEN_1 and _2 continue working
Provider keys¶
Configure only the providers you use — others can be omitted. At least one provider must be set, or all requests return 503.
| Variable | Provider | Key format |
|---|---|---|
ANTHROPIC_API_KEY |
Anthropic Claude | sk-ant-… |
OPENAI_API_KEY |
OpenAI | sk-proj-… |
GROQ_API_KEY |
Groq | gsk_… |
GEMINI_API_KEY |
Google Gemini | AIza… |
MOONSHOT_API_KEY |
Moonshot AI | sk-… |
DEEPSEEK_API_KEY |
DeepSeek | sk-… |
MISTRAL_API_KEY |
Mistral AI | — |
XAI_API_KEY |
xAI Grok | — |
TOGETHER_API_KEY |
Together AI | — |
FIREWORKS_API_KEY |
Fireworks AI | — |
CEREBRAS_API_KEY |
Cerebras | — |
Local providers (Ollama, LM Studio, vLLM) need no API key — they are keyless.
Multi-key round-robin¶
Add numbered variants for any provider to distribute load across multiple keys. The proxy cycles through them in order:
ANTHROPIC_API_KEY=sk-ant-key-a
ANTHROPIC_API_KEY_1=sk-ant-key-b
ANTHROPIC_API_KEY_2=sk-ant-key-c
# Requests rotate: key-a → key-b → key-c → key-a → …
This applies to all cloud providers. The GET /health response shows how many keys are active per provider.
Failover chain¶
POST /v1/chat/completions — the unified endpoint — automatically tries providers in order if one times out or rate-limits:
Format: provider:model,provider:model,… — each provider must have an API key configured. Unknown providers are skipped with a warning at startup. Failover triggers on timeout (30 s) or HTTP 429.
See Provider routing for the full unified endpoint reference.
Observability¶
| Variable | Default | Description |
|---|---|---|
AUDIT_LOG_PATH |
/var/log/keybridge/audit.jsonl |
Append-only JSON-lines audit log. Each line is one request. The API key is SHA-256 hashed — plaintext keys are never written. |
METRICS_TOKEN |
(unset) | When set, GET /health and GET /metrics require a matching Authorization: Bearer <token> or x-api-key: <token> header. Leave unset to keep them public (default). |
The GET /metrics endpoint returns live in-memory stats: request count, token totals, p50/p99 latency, and error counts. Stats reset on process restart.
Securing observability endpoints
In production, set METRICS_TOKEN to a random secret so only your monitoring system can reach /health and /metrics. This prevents competitor traffic analysis via your public health endpoint.
Rate limiting¶
| Variable | Default | Description |
|---|---|---|
PROXY_TOKEN_RPM |
600 |
Requests per minute allowed per token (sliding-window, in-memory). 600 = 10 req/s. Set to 0 to disable. Note: not shared across multiple worker processes — for multi-worker deployments use an external rate limiter. |
mTLS (mutual TLS)¶
keybridge supports mTLS as an alternative or complement to token-based auth:
| Variable | Description |
|---|---|
SSL_CERTFILE |
Path to the server's TLS certificate (PEM) |
SSL_KEYFILE |
Path to the server's TLS private key (PEM) |
SSL_KEYFILE_PASSWORD |
Passphrase for an encrypted key file (optional) |
MTLS_CA_CERT |
CA certificate (PEM) that must have signed the client cert. When set, clients must present a valid cert or the TLS handshake is rejected before any HTTP code runs. |
Auth modes based on env vars set:
| Vars set | Auth mode |
|---|---|
PROXY_TOKEN only |
Token auth (default) |
SSL_* + MTLS_CA_CERT |
mTLS only (no token needed) |
PROXY_TOKEN + SSL_* + MTLS_CA_CERT |
Both — defense in depth |
Quick CA + cert setup:
# CA
openssl genrsa -out ca.key 4096
openssl req -x509 -new -key ca.key -days 3650 -out ca.crt -subj "/CN=keybridge-ca"
# Server cert
openssl genrsa -out server.key 4096
openssl req -new -key server.key -out server.csr -subj "/CN=keybridge"
openssl x509 -req -in server.csr -CA ca.crt -CAkey ca.key -CAcreateserial -days 365 -out server.crt
# Client cert (for mTLS)
openssl genrsa -out client.key 4096
openssl req -new -key client.key -out client.csr -subj "/CN=client"
openssl x509 -req -in client.csr -CA ca.crt -CAkey ca.key -CAcreateserial -days 365 -out client.crt
# Run with mTLS
docker run -d -p 8443:8443 \
-e SSL_CERTFILE=/certs/server.crt -e SSL_KEYFILE=/certs/server.key \
-e MTLS_CA_CERT=/certs/ca.crt \
-v /path/to/certs:/certs \
ghcr.io/iagop03/keybridge:latest
Limits¶
| Variable | Default | Description |
|---|---|---|
MAX_BODY_BYTES |
10485760 (10 MB) |
Maximum request body size. Larger requests return 413. |
MAX_CONCURRENT |
20 |
Default max concurrent upstream requests per provider. |
MAX_CONCURRENT_ANTHROPIC |
20 |
Per-provider override — use this to set tighter limits on expensive providers. Replace ANTHROPIC with any provider name in uppercase. |
Upstream URL overrides¶
Security — SSRF risk
Each *_BASE_URL variable is used verbatim as the upstream target. An attacker who can write to these variables can redirect proxy traffic to arbitrary hosts, including internal services. Treat them with the same care as API keys: set them at container build time or via a secrets manager, never accept them from user input.
| Variable | Default |
|---|---|
ANTHROPIC_BASE_URL |
https://api.anthropic.com |
OPENAI_BASE_URL |
https://api.openai.com |
GROQ_BASE_URL |
https://api.groq.com/openai |
GEMINI_BASE_URL |
https://generativelanguage.googleapis.com/v1beta/openai |
MOONSHOT_BASE_URL |
https://api.moonshot.cn/v1 |
DEEPSEEK_BASE_URL |
https://api.deepseek.com/v1 |
MISTRAL_BASE_URL |
https://api.mistral.ai/v1 |
XAI_BASE_URL |
https://api.x.ai/v1 |
TOGETHER_BASE_URL |
https://api.together.xyz/v1 |
FIREWORKS_BASE_URL |
https://api.fireworks.ai/inference/v1 |
CEREBRAS_BASE_URL |
https://api.cerebras.ai/v1 |
OLLAMA_BASE_URL |
http://localhost:11434 |
LMSTUDIO_BASE_URL |
http://localhost:1234/v1 |
VLLM_BASE_URL |
http://localhost:8000/v1 |
Server¶
| Variable | Default | Description |
|---|---|---|
HOST |
0.0.0.0 |
Bind address |
PORT |
8080 |
Bind port. Port 8080 is in Cloudflare's supported HTTP list. |
Docker¶
services:
proxy:
image: ghcr.io/iagop03/keybridge:latest
ports:
- "8080:8080"
environment:
PROXY_TOKEN: ${PROXY_TOKEN}
ANTHROPIC_API_KEY: ${ANTHROPIC_API_KEY}
OPENAI_API_KEY: ${OPENAI_API_KEY}
FAILOVER_CHAIN: "anthropic:claude-opus-5,openai:gpt-4o"
AUDIT_LOG_PATH: "/var/log/keybridge/audit.jsonl"
MAX_CONCURRENT: "20"
# add only the providers you use
volumes:
- audit_logs:/var/log/keybridge
volumes:
audit_logs:
OpenTelemetry tracing¶
keybridge can export distributed traces to any OTLP-compatible collector (Jaeger, Tempo, Honeycomb, Datadog, etc.).
Install the optional dependency and set the endpoint:
| Variable | Default | Description |
|---|---|---|
OTLP_ENDPOINT |
(unset) | gRPC endpoint for the OTLP exporter. When unset, tracing is silently disabled — no error, no overhead. |
Trace spans cover: each inbound request, upstream provider call, auth check, and audit log write. FastAPI and httpx are auto-instrumented.
Azure OpenAI¶
Set AZURE_OPENAI_ENDPOINT (e.g. https://myresource.openai.azure.com) and AZURE_OPENAI_API_KEY. Route requests to /azure/openai/deployments/{deployment}/…. The proxy injects the key as the api-key header (not Authorization: Bearer).