LLM Proxy Bridge¶
The one path every LLM call takes β which is what makes routing, fallback and token accounting possible at all.
What it is¶
An HTTP service on the host at port 12435. Containers cannot reach host credentials or CLI tools, so without it every containerised workload would fall through to paid API providers. The bridge lets them use subscription accounts instead, at no marginal cost.
The endpoint¶
Not the OpenAI-shaped /v1/chat/completions. Body is { process, messages, complexity? }; the response is { content, provider, model, tokens, latencyMs } β not OpenAI-wrapped, and provider is the account id.
The port that trips everyone¶
12435 is the proxy. 3033 is the Health API. Posting a completion to 3033 returns Cannot POST /api/complete β a bare 404 that names nothing.
Health¶
After changing its source¶
Its dist/ is gitignored and the runtime imports the compiled output, so edits to src/ do not reach the running daemon until you rebuild and restart it.
Why it exists¶
Inside a container there are no host credentials and no host CLIs. Without a bridge, the provider chain in every containerised workload falls through to metered APIs β which is both expensive and a different set of models than the interactive session is using.

Because everything goes through one place, three things become possible that otherwise are not: routing decisions can be made centrally and reproducibly, fallback can be described as configuration rather than scattered code, and every call can be accounted for.
Talking to it¶
curl -s localhost:12435/api/complete \
-H 'content-type: application/json' \
-d '{"process":"my-service","messages":[{"role":"user","content":"say OK"}],"complexity":"small"}'
Three things about that request are worth stating explicitly, because each has caused real confusion:
The path is /api/complete. Reaching for the OpenAI shape gets a 404.
process is what makes accounting work. It names the calling service and is what the token dashboard attributes by. A call without one lands under unknown.
complexity is the band, and it is the field that controls cost. A taskType field does nothing β it is read by nothing in the proxy, which is how a background classifier once ran on an expensive model while appearing to declare itself cheap.
Two ports, one common mistake¶
12435 is this proxy. 3033 is the Health API. They are unrelated services, and posting a completion to the wrong one produces Cannot POST /api/complete rather than anything that names the problem.
Subscription accounts and their fallbacks¶
The bridge serves subscription providers first. For the Claude path there are two tiers: a fast direct OAuth call, and a slower CLI fallback used when the direct path is rejected.
The fallback is not merely slower β it routes through a different rate-limit bucket on the same subscription, which is why it can succeed when the direct path is being throttled. It also carries a large auto-injected system prompt, billed as cache creation, so its token figures look very different for the same work.
Rebuilding after a change¶
The compiled dist/ is gitignored and the running service imports it, so a fresh clone cannot start the proxy and an edit to src/ has no effect until it is rebuilt and the daemon restarted. Configuration is different β the YAML files hot-reload on save, and the bridge's own .mjs files do not.
Checking it¶
curl -s localhost:12435/health | jq '.networkMode, .egress'
curl -s 'localhost:12435/api/llm/routing/resolve?job=<job>' | jq -r .summary
The first answers "can this process reach the internet, and on what network"; the second answers "where would this call go". They fail independently, and a routing decision can be perfectly correct while egress is broken.
Extracted to standalone package: The LLM layer is now provided by
@rapid/llm-proxy. The localsrc/llm-proxy/llm-proxy.mjsis a thin wrapper that delegates to the package.
Overview¶
The LLM Proxy Bridge enables Docker containers to access host-side LLM capabilities. It runs on the host machine (port 12435) and serves as an HTTP intermediary between containerised workloads and the subscription LLM providers:
- Copilot β direct HTTP POST to the Copilot API using OAuth tokens from
~/.local/share/opencode/auth.json. Parallelism-optimised (~0.77s effective per call at 10 concurrent). - Claude Code (Max subscription) β a two-tier dispatch (since 2026-05-19):
- Direct OAuth path (fast): POST to
api.anthropic.com/v1/messageswithAuthorization: Bearer <oauth>read from the macOS keychain. ~0.9s, real token counts (9 in / 4 out forsay OK). - CLI fallback (slower): when the direct path returns 401/403/429, the dispatcher falls back to spawning the
claude -psubprocess. The CLI auto-injects ~16-22K tokens of system prompt (billed ascache_creation) and takes ~10-14s per call, but routes through a different Anthropic rate-limit bucket on the same Max subscription β empirically sonnet/opus succeed via CLI while the bearer endpoint 429s.
- Direct OAuth path (fast): POST to
Port: 12435 (host)
Why it exists: Inside Docker, host-side credentials and CLI tools are unavailable. Without the proxy, the LLM provider chain falls back to paid API providers (Groq, Anthropic, OpenAI). The proxy bridges this gap, letting Docker workloads use subscription-based providers at zero incremental cost.
Architecture¶

sequenceDiagram
participant D as Docker Container
participant P as LLM Proxy Bridge (Host:12435)
participant ApiAnt as api.anthropic.com
participant Cli as claude CLI
participant ApiCop as Copilot API
D->>P: POST /api/complete {provider, messages, process}
alt provider = copilot
P->>ApiCop: POST /chat/completions (OAuth bearer)
ApiCop-->>P: JSON response (~2s)
else provider = claude-code (Path 1: direct OAuth)
P->>ApiAnt: POST /v1/messages (Bearer <oauth>)
alt 200 OK
ApiAnt-->>P: completion (~0.9s, real tokens)
else 401/403/429
ApiAnt-->>P: error
P->>Cli: claude -p --model <m> --tools '' (Path 2 fallback)
Cli-->>P: completion (~10-14s, includes cache_creation)
end
end
P-->>D: {content, tokens, latencyMs} Provider routing¶
This section previously described an auto-route that no longer exists. It said preference depended on the caller's session type (
claude-codefirst for Claude sessions,copilotfirst for OpenCode/corporate), and that operator pins persisted to.data/llm-proxy/llm-settings.json. That was theprocessOverridesruntime-state mechanism, and the session-type ternary insideserver.mjsβ both are deleted. Routing is now config-driven end to end.
Every decision comes from two version-controlled YAML files in the proxy repo, and nothing else: config/llm-routing.yaml (which provider + model) and config/llm-fallback.yaml (chains + guards). There is no auto-route, no session-type heuristic, and no runtime pin file. Lookup is three fixed steps β routes["<job>/<agent>"] β routes["<job>"] β defaults[<class>] β and the complexity band picks the model from the chosen provider's own table, so a fallback always lands on something that provider actually serves.
Providers are named by the account that gets billed (claude-code-max, gh-copilot), not the company. The old company-level labels above (claude-code, copilot, anthropic) survive only as impl values in the config and as historical token_usage rows the dashboard maps forward.
Edit through Token Usage β Settings β the Routing and Fallback tabs write the YAML back with comments preserved and validation before write, and the Flow tab draws the result as a graph weighted by real traffic.

Within claude-code-max, the directβCLI fallback ladder still applies, because both paths run on the same Max subscription.
Full reference: LLM Routing. Prefer it over this section β it is maintained alongside the code that obeys it.
Observed behaviour¶
A snapshot of the Token Usage page's Recent Calls table once the direct-OAuth path went live. The claude-code/claude-haiku-4.5 rows show the new 9 input / 4 output / ~1s envelope from the direct path; the older claude-code/claude-haiku-4.5 rows above (with 14.9K input tokens / 7+s latency) are pre-fix CLI calls retained for comparison.

Claude CLI Worker Pool (v7.3)¶
The Worker Pool is the v7.3 evolution of the claude-code CLI-fallback path. Where the legacy fallback spawned a cold claude -p subprocess per request β paying the CLI's full ~16-22K-token system-prompt warm-up on every call (~10-14s) β the pool keeps a claude -p subprocess warm and reused, cutting steady-state fallback latency from ~14s to ~2-3s.
It lives in the standalone proxy bridge (proxy-bridge/worker-pool.mjs, wired into proxy-bridge/server.mjs's claude-code dispatcher). It only affects the claude-code provider β copilot remains a direct HTTP POST and is never pooled.

Two-tier dispatch¶
When the direct-OAuth fast path 401/403/429s and the dispatcher falls back to the CLI, the fallback itself is now two-tiered:
- Tier 1 β warm pool (fast path):
WorkerPool.complete(body, abortSignal, overflowFn)keys requests bymodel :: sha256(systemPrompt)[:16]. The first request for a given key lazily spawns a persistentclaude -p --input-format stream-json --output-format stream-jsonsubprocess (concurrency-1, FIFO queue) and reuses it warm on subsequent calls. ~2-3s steady-state. - Tier 2 β overflow (cold one-shot):
completeClaudeCodeViaCLIruns a cold one-shotexecFileofclaude -pper request (~10-14s, the legacy behaviour). The pool falls through to overflow when all workers for a key are busy, when the key is in crash-cooldown, or when the worker pool is disabled viaLLM_PROXY_DISABLE_WORKER_POOL=1(the GUARD-01 escape hatch).
Lifecycle¶
The pool guarantees four lifecycle behaviours (WLIFE-01..04):

- WLIFE-01 β lazy spawn: zero workers exist at boot. The first claude-code fallback request spawns exactly one worker; idle keys cost nothing.
- WLIFE-02 β idle eviction: a worker idle past
LLM_PROXY_WORKER_IDLE_MS(default 30 min) disposes itself via an unref'd timer and is dropped, freeing RAM in quiet periods. The next same-key request lazily respawns it. - WLIFE-03 β crash recovery: a worker that crashes or EPIPEs mid-request surfaces its in-flight and queued jobs as RETRYABLE (never a hang). Per-key crash tracking puts a storming key into cooldown β routing it to overflow β once it crosses
LLM_PROXY_WORKER_CRASH_THRESHOLDcrashes withinLLM_PROXY_WORKER_CRASH_WINDOW_MS, preventing a spawnβcrashβrespawn storm. - WLIFE-04 β cancellation: a client disconnect/abort SIGTERMs and disposes the in-flight worker and drops it synchronously, so the next same-key request gets a fresh cold worker. A queued (not-yet-running) abort only dequeues that one job, leaving the live worker untouched.
Workers also recycle proactively after LLM_PROXY_WORKER_MAX_REQUESTS requests or LLM_PROXY_WORKER_MAX_INPUT_TOKENS cumulative input tokens, and an LRU prompt-cap (LLM_PROXY_WORKER_PROMPT_CAP) bounds the number of distinct (model Γ prompt) pools kept alive. See Worker Pool tuning for the full env-knob reference.
API Endpoints¶
GET /health¶
Returns proxy status and provider availability.
Response:
{
"status": "ok",
"providers": {
"copilot": {
"available": true,
"mode": "direct-http",
"lastChecked": 1708600000000
},
"claude-code": {
"available": true,
"version": "2.1.50",
"lastChecked": 1708600000000
}
},
"uptime": 3600,
"inFlightRequests": 0
}
POST /api/complete¶
Forward a completion request to an LLM provider.
Request:
{
"provider": "copilot",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain dependency injection."}
],
"model": "claude-sonnet-4.5",
"maxTokens": 1000,
"temperature": 0.5,
"tier": "standard"
}
Response (200):
{
"content": "Dependency injection is...",
"provider": "copilot",
"model": "claude-sonnet-4.5",
"tokens": {"input": 25, "output": 150, "total": 175},
"latencyMs": 2100
}
Error Responses:
| Status | Type | Meaning |
|---|---|---|
| 400 | VALIDATION_ERROR | Missing required fields or unknown provider |
| 401 | AUTH_ERROR | OAuth token expired or missing. For claude-code, the dispatcher will normally have already fallen back to the CLI before surfacing this. |
| 429 | QUOTA_EXHAUSTED | Provider rate limit reached. For claude-code this surfaces only when both the direct OAuth bearer AND the CLI fallback return 429 β a true hard limit on the Max subscription. The common case (sonnet via bearer 429, CLI succeeds) is invisible to the caller. |
| 503 | PROVIDER_UNAVAILABLE | Provider not configured or unreachable |
| 504 | TIMEOUT | Request timed out |
| 500 | PROVIDER_ERROR | Upstream error |
Environment escape hatches (claude-code)¶
| Env var | Effect |
|---|---|
LLM_PROXY_DISABLE_CLAUDE_DIRECT=1 | Force every claude-code call through the legacy CLI path (skip the direct OAuth fast path). Useful if Anthropic tightens the bearer-endpoint allowlist or for debugging. |
Configuration¶
Environment Variables¶
| Variable | Default | Description |
|---|---|---|
LLM_CLI_PROXY_PORT | 12435 | Port the proxy listens on |
LLM_CLI_PROXY_URL | - | Set in Docker containers to http://host.docker.internal:12435 |
Docker Compose¶
The Docker container connects via host.docker.internal:
Quick Start¶
# Start the proxy bridge (thin wrapper delegates to @rapid/llm-proxy)
node src/llm-proxy/llm-proxy.mjs
# Or use the standalone package directly
npx @rapid/llm-proxy
Auto-Start¶
The proxy starts automatically when launching coding (via bin/coding). The startup script (scripts/start-services-robust.js) handles:
- Checking if already running on port 12435
- Spawning the process with retry logic
- Health check via
GET /health
Health Monitoring¶
The proxy is monitored by the health verification system:
- Severity: Warning (optional service)
- Check type: HTTP health (
GET /health) - Auto-heal: Disabled (host-side service, not Docker-managed)
- Dashboard: Appears in System Health Dashboard service checks
Troubleshooting¶
Proxy not starting¶
- Check if port is already in use:
lsof -ti:12435 - Check logs:
node src/llm-proxy/llm-proxy.mjs
Copilot auth failure¶
OAuth tokens are read from ~/.local/share/opencode/auth.json. If expired:
Claude CLI not found¶
The CLI fallback path requires claude on the host:
If the direct OAuth path works, the CLI is only needed for sonnet/opus calls (rate-limited on the bearer endpoint). Without the CLI, those calls will surface as QUOTA_EXHAUSTED.
claude-code calls take ~14s instead of ~1s¶
The dispatcher fell back to the CLI path. Check the proxy log:
Look for direct API rate-limited ... falling back to CLI. This is expected today for sonnet/opus β Anthropic's OAuth bearer endpoint per-model rate limit. The CLI fallback uses the same Max subscription but a different rate-limit bucket. As of v7.3 the Claude CLI Worker Pool keeps the fallback subprocess warm, bringing steady-state fallback latency down to ~2-3s; cold one-shot overflow calls still take ~10-14s.
"say OK" health-coordinator probe burns 16K tokens per call¶
This is the symptom of health-coordinator being routed through the CLI path. Verify the direct OAuth path is enabled:
If set to 1, unset it and restart the proxy:
Healthy state: say OK calls report 9 input + 4 output tokens at ~0.9s.
Health dashboard shows degraded¶
Ensure the health check endpoint uses localhost:12435 (not host.docker.internal, which only resolves inside Docker).
Full Documentation¶
See the @rapid/llm-proxy package documentation for:
- Architecture β provider stack, tier routing, circuit breaker
- Providers β per-provider setup and auth
- Proxy Bridge β Docker bridge details
- Configuration β YAML config reference
Related¶
- LLM Architecture β Unified LLM provider layer
- LLM Providers Guide β Provider configuration