Skip to content

Configuration

Environment variables

Required

Variable Description
PROXY_TOKEN UUID token generated by the platform in Settings → LLM Mode → Proxy. Validated on every request with constant-time comparison.

Multi-token auth (hot rotation)

To accept multiple tokens simultaneously — useful when rotating keys without downtime — add numbered variants. Any listed token is accepted:

PROXY_TOKEN=token-primary
PROXY_TOKEN_1=token-secondary   # add new token
PROXY_TOKEN_2=token-backup
# Remove PROXY_TOKEN later; PROXY_TOKEN_1 and _2 continue working

Provider keys

Configure only the providers you use — others can be omitted. At least one provider must be set, or all requests return 503.

Variable Provider Key format
ANTHROPIC_API_KEY Anthropic Claude sk-ant-…
OPENAI_API_KEY OpenAI sk-proj-…
GROQ_API_KEY Groq gsk_…
GEMINI_API_KEY Google Gemini AIza…
MOONSHOT_API_KEY Moonshot AI sk-…
DEEPSEEK_API_KEY DeepSeek sk-…
MISTRAL_API_KEY Mistral AI —
XAI_API_KEY xAI Grok —
TOGETHER_API_KEY Together AI —
FIREWORKS_API_KEY Fireworks AI —
CEREBRAS_API_KEY Cerebras —

Local providers (Ollama, LM Studio, vLLM) need no API key — they are keyless.

Multi-key round-robin

Add numbered variants for any provider to distribute load across multiple keys. The proxy cycles through them in order:

ANTHROPIC_API_KEY=sk-ant-key-a
ANTHROPIC_API_KEY_1=sk-ant-key-b
ANTHROPIC_API_KEY_2=sk-ant-key-c
# Requests rotate: key-a → key-b → key-c → key-a → …

This applies to all cloud providers. The GET /health response shows how many keys are active per provider.

Failover chain

POST /v1/chat/completions — the unified endpoint — automatically tries providers in order if one times out or rate-limits:

FAILOVER_CHAIN=anthropic:claude-opus-5,openai:gpt-4o,groq:llama-3.3-70b-versatile

Format: provider:model,provider:model,… — each provider must have an API key configured. Unknown providers are skipped with a warning at startup. Failover triggers on timeout (30 s) or HTTP 429.

See Provider routing for the full unified endpoint reference.

Observability

Variable Default Description
AUDIT_LOG_PATH /var/log/keybridge/audit.jsonl Append-only JSON-lines audit log. Each line is one request. The API key is SHA-256 hashed — plaintext keys are never written.
METRICS_TOKEN (unset) When set, GET /health and GET /metrics require a matching Authorization: Bearer <token> or x-api-key: <token> header. Leave unset to keep them public (default).

The GET /metrics endpoint returns live in-memory stats: request count, token totals, p50/p99 latency, and error counts. Stats reset on process restart.

Securing observability endpoints

In production, set METRICS_TOKEN to a random secret so only your monitoring system can reach /health and /metrics. This prevents competitor traffic analysis via your public health endpoint.

Rate limiting

Variable Default Description
PROXY_TOKEN_RPM 600 Requests per minute allowed per token (sliding-window, in-memory). 600 = 10 req/s. Set to 0 to disable. Note: not shared across multiple worker processes — for multi-worker deployments use an external rate limiter.

mTLS (mutual TLS)

keybridge supports mTLS as an alternative or complement to token-based auth:

Variable Description
SSL_CERTFILE Path to the server's TLS certificate (PEM)
SSL_KEYFILE Path to the server's TLS private key (PEM)
SSL_KEYFILE_PASSWORD Passphrase for an encrypted key file (optional)
MTLS_CA_CERT CA certificate (PEM) that must have signed the client cert. When set, clients must present a valid cert or the TLS handshake is rejected before any HTTP code runs.

Auth modes based on env vars set:

Vars set Auth mode
PROXY_TOKEN only Token auth (default)
SSL_* + MTLS_CA_CERT mTLS only (no token needed)
PROXY_TOKEN + SSL_* + MTLS_CA_CERT Both — defense in depth

Quick CA + cert setup:

# CA
openssl genrsa -out ca.key 4096
openssl req -x509 -new -key ca.key -days 3650 -out ca.crt -subj "/CN=keybridge-ca"
# Server cert
openssl genrsa -out server.key 4096
openssl req -new -key server.key -out server.csr -subj "/CN=keybridge"
openssl x509 -req -in server.csr -CA ca.crt -CAkey ca.key -CAcreateserial -days 365 -out server.crt
# Client cert (for mTLS)
openssl genrsa -out client.key 4096
openssl req -new -key client.key -out client.csr -subj "/CN=client"
openssl x509 -req -in client.csr -CA ca.crt -CAkey ca.key -CAcreateserial -days 365 -out client.crt

# Run with mTLS
docker run -d -p 8443:8443 \
  -e SSL_CERTFILE=/certs/server.crt -e SSL_KEYFILE=/certs/server.key \
  -e MTLS_CA_CERT=/certs/ca.crt \
  -v /path/to/certs:/certs \
  ghcr.io/iagop03/keybridge:latest

Limits

Variable Default Description
MAX_BODY_BYTES 10485760 (10 MB) Maximum request body size. Larger requests return 413.
MAX_CONCURRENT 20 Default max concurrent upstream requests per provider.
MAX_CONCURRENT_ANTHROPIC 20 Per-provider override — use this to set tighter limits on expensive providers. Replace ANTHROPIC with any provider name in uppercase.

Upstream URL overrides

Security — SSRF risk

Each *_BASE_URL variable is used verbatim as the upstream target. An attacker who can write to these variables can redirect proxy traffic to arbitrary hosts, including internal services. Treat them with the same care as API keys: set them at container build time or via a secrets manager, never accept them from user input.

Variable Default
ANTHROPIC_BASE_URL https://api.anthropic.com
OPENAI_BASE_URL https://api.openai.com
GROQ_BASE_URL https://api.groq.com/openai
GEMINI_BASE_URL https://generativelanguage.googleapis.com/v1beta/openai
MOONSHOT_BASE_URL https://api.moonshot.cn/v1
DEEPSEEK_BASE_URL https://api.deepseek.com/v1
MISTRAL_BASE_URL https://api.mistral.ai/v1
XAI_BASE_URL https://api.x.ai/v1
TOGETHER_BASE_URL https://api.together.xyz/v1
FIREWORKS_BASE_URL https://api.fireworks.ai/inference/v1
CEREBRAS_BASE_URL https://api.cerebras.ai/v1
OLLAMA_BASE_URL http://localhost:11434
LMSTUDIO_BASE_URL http://localhost:1234/v1
VLLM_BASE_URL http://localhost:8000/v1

Server

Variable Default Description
HOST 0.0.0.0 Bind address
PORT 8080 Bind port. Port 8080 is in Cloudflare's supported HTTP list.

Docker

services:
  proxy:
    image: ghcr.io/iagop03/keybridge:latest
    ports:
      - "8080:8080"
    environment:
      PROXY_TOKEN: ${PROXY_TOKEN}
      ANTHROPIC_API_KEY: ${ANTHROPIC_API_KEY}
      OPENAI_API_KEY: ${OPENAI_API_KEY}
      FAILOVER_CHAIN: "anthropic:claude-opus-5,openai:gpt-4o"
      AUDIT_LOG_PATH: "/var/log/keybridge/audit.jsonl"
      MAX_CONCURRENT: "20"
      # add only the providers you use
    volumes:
      - audit_logs:/var/log/keybridge

volumes:
  audit_logs:

OpenTelemetry tracing

keybridge can export distributed traces to any OTLP-compatible collector (Jaeger, Tempo, Honeycomb, Datadog, etc.).

Install the optional dependency and set the endpoint:

pip install "keybridge[otel]"
OTLP_ENDPOINT=http://otel-collector:4317 keybridge
Variable Default Description
OTLP_ENDPOINT (unset) gRPC endpoint for the OTLP exporter. When unset, tracing is silently disabled — no error, no overhead.

Trace spans cover: each inbound request, upstream provider call, auth check, and audit log write. FastAPI and httpx are auto-instrumented.

Azure OpenAI

Set AZURE_OPENAI_ENDPOINT (e.g. https://myresource.openai.azure.com) and AZURE_OPENAI_API_KEY. Route requests to /azure/openai/deployments/{deployment}/…. The proxy injects the key as the api-key header (not Authorization: Bearer).

AZURE_OPENAI_ENDPOINT=https://myresource.openai.azure.com
AZURE_OPENAI_API_KEY=your-azure-key