Skip to content

LLM Proxy Bridge

The one path every LLM call takes β€” which is what makes routing, fallback and token accounting possible at all.

What it is

An HTTP service on the host at port 12435. Containers cannot reach host credentials or CLI tools, so without it every containerised workload would fall through to paid API providers. The bridge lets them use subscription accounts instead, at no marginal cost.

The endpoint

POST http://localhost:12435/api/complete

Not the OpenAI-shaped /v1/chat/completions. Body is { process, messages, complexity? }; the response is { content, provider, model, tokens, latencyMs } β€” not OpenAI-wrapped, and provider is the account id.

The port that trips everyone

12435 is the proxy. 3033 is the Health API. Posting a completion to 3033 returns Cannot POST /api/complete β€” a bare 404 that names nothing.

Health

curl -s localhost:12435/health | jq '.networkMode, .egress'

After changing its source

Its dist/ is gitignored and the runtime imports the compiled output, so edits to src/ do not reach the running daemon until you rebuild and restart it.

Why it exists

Inside a container there are no host credentials and no host CLIs. Without a bridge, the provider chain in every containerised workload falls through to metered APIs β€” which is both expensive and a different set of models than the interactive session is using.

LLM Proxy Bridge Architecture

Because everything goes through one place, three things become possible that otherwise are not: routing decisions can be made centrally and reproducibly, fallback can be described as configuration rather than scattered code, and every call can be accounted for.

Talking to it

curl -s localhost:12435/api/complete \
  -H 'content-type: application/json' \
  -d '{"process":"my-service","messages":[{"role":"user","content":"say OK"}],"complexity":"small"}'

Three things about that request are worth stating explicitly, because each has caused real confusion:

The path is /api/complete. Reaching for the OpenAI shape gets a 404.

process is what makes accounting work. It names the calling service and is what the token dashboard attributes by. A call without one lands under unknown.

complexity is the band, and it is the field that controls cost. A taskType field does nothing β€” it is read by nothing in the proxy, which is how a background classifier once ran on an expensive model while appearing to declare itself cheap.

Two ports, one common mistake

12435 is this proxy. 3033 is the Health API. They are unrelated services, and posting a completion to the wrong one produces Cannot POST /api/complete rather than anything that names the problem.

Subscription accounts and their fallbacks

The bridge serves subscription providers first. For the Claude path there are two tiers: a fast direct OAuth call, and a slower CLI fallback used when the direct path is rejected.

The fallback is not merely slower β€” it routes through a different rate-limit bucket on the same subscription, which is why it can succeed when the direct path is being throttled. It also carries a large auto-injected system prompt, billed as cache creation, so its token figures look very different for the same work.

Rebuilding after a change

The compiled dist/ is gitignored and the running service imports it, so a fresh clone cannot start the proxy and an edit to src/ has no effect until it is rebuilt and the daemon restarted. Configuration is different β€” the YAML files hot-reload on save, and the bridge's own .mjs files do not.

Checking it

curl -s localhost:12435/health | jq '.networkMode, .egress'
curl -s 'localhost:12435/api/llm/routing/resolve?job=<job>' | jq -r .summary

The first answers "can this process reach the internet, and on what network"; the second answers "where would this call go". They fail independently, and a routing decision can be perfectly correct while egress is broken.

Extracted to standalone package: The LLM layer is now provided by @rapid/llm-proxy. The local src/llm-proxy/llm-proxy.mjs is a thin wrapper that delegates to the package.

Overview

The LLM Proxy Bridge enables Docker containers to access host-side LLM capabilities. It runs on the host machine (port 12435) and serves as an HTTP intermediary between containerised workloads and the subscription LLM providers:

  • Copilot β€” direct HTTP POST to the Copilot API using OAuth tokens from ~/.local/share/opencode/auth.json. Parallelism-optimised (~0.77s effective per call at 10 concurrent).
  • Claude Code (Max subscription) β€” a two-tier dispatch (since 2026-05-19):
    1. Direct OAuth path (fast): POST to api.anthropic.com/v1/messages with Authorization: Bearer <oauth> read from the macOS keychain. ~0.9s, real token counts (9 in / 4 out for say OK).
    2. CLI fallback (slower): when the direct path returns 401/403/429, the dispatcher falls back to spawning the claude -p subprocess. The CLI auto-injects ~16-22K tokens of system prompt (billed as cache_creation) and takes ~10-14s per call, but routes through a different Anthropic rate-limit bucket on the same Max subscription β€” empirically sonnet/opus succeed via CLI while the bearer endpoint 429s.

Port: 12435 (host)

Why it exists: Inside Docker, host-side credentials and CLI tools are unavailable. Without the proxy, the LLM provider chain falls back to paid API providers (Groq, Anthropic, OpenAI). The proxy bridges this gap, letting Docker workloads use subscription-based providers at zero incremental cost.


Architecture

LLM Proxy Bridge Architecture

sequenceDiagram
    participant D as Docker Container
    participant P as LLM Proxy Bridge (Host:12435)
    participant ApiAnt as api.anthropic.com
    participant Cli as claude CLI
    participant ApiCop as Copilot API

    D->>P: POST /api/complete {provider, messages, process}
    alt provider = copilot
        P->>ApiCop: POST /chat/completions (OAuth bearer)
        ApiCop-->>P: JSON response (~2s)
    else provider = claude-code (Path 1: direct OAuth)
        P->>ApiAnt: POST /v1/messages (Bearer <oauth>)
        alt 200 OK
            ApiAnt-->>P: completion (~0.9s, real tokens)
        else 401/403/429
            ApiAnt-->>P: error
            P->>Cli: claude -p --model <m> --tools '' (Path 2 fallback)
            Cli-->>P: completion (~10-14s, includes cache_creation)
        end
    end
    P-->>D: {content, tokens, latencyMs}

Provider routing

This section previously described an auto-route that no longer exists. It said preference depended on the caller's session type (claude-code first for Claude sessions, copilot first for OpenCode/corporate), and that operator pins persisted to .data/llm-proxy/llm-settings.json. That was the processOverrides runtime-state mechanism, and the session-type ternary inside server.mjs β€” both are deleted. Routing is now config-driven end to end.

Every decision comes from two version-controlled YAML files in the proxy repo, and nothing else: config/llm-routing.yaml (which provider + model) and config/llm-fallback.yaml (chains + guards). There is no auto-route, no session-type heuristic, and no runtime pin file. Lookup is three fixed steps β€” routes["<job>/<agent>"] β†’ routes["<job>"] β†’ defaults[<class>] β€” and the complexity band picks the model from the chosen provider's own table, so a fallback always lands on something that provider actually serves.

Providers are named by the account that gets billed (claude-code-max, gh-copilot), not the company. The old company-level labels above (claude-code, copilot, anthropic) survive only as impl values in the config and as historical token_usage rows the dashboard maps forward.

Edit through Token Usage β†’ Settings β€” the Routing and Fallback tabs write the YAML back with comments preserved and validation before write, and the Flow tab draws the result as a graph weighted by real traffic.

LLM Routing settings β€” Routing tab

Within claude-code-max, the direct→CLI fallback ladder still applies, because both paths run on the same Max subscription.

Full reference: LLM Routing. Prefer it over this section β€” it is maintained alongside the code that obeys it.

Observed behaviour

A snapshot of the Token Usage page's Recent Calls table once the direct-OAuth path went live. The claude-code/claude-haiku-4.5 rows show the new 9 input / 4 output / ~1s envelope from the direct path; the older claude-code/claude-haiku-4.5 rows above (with 14.9K input tokens / 7+s latency) are pre-fix CLI calls retained for comparison.

Recent LLM calls β€” direct-OAuth vs. legacy CLI path


Claude CLI Worker Pool (v7.3)

The Worker Pool is the v7.3 evolution of the claude-code CLI-fallback path. Where the legacy fallback spawned a cold claude -p subprocess per request β€” paying the CLI's full ~16-22K-token system-prompt warm-up on every call (~10-14s) β€” the pool keeps a claude -p subprocess warm and reused, cutting steady-state fallback latency from ~14s to ~2-3s.

It lives in the standalone proxy bridge (proxy-bridge/worker-pool.mjs, wired into proxy-bridge/server.mjs's claude-code dispatcher). It only affects the claude-code provider β€” copilot remains a direct HTTP POST and is never pooled.

Claude CLI Worker Pool architecture

Two-tier dispatch

When the direct-OAuth fast path 401/403/429s and the dispatcher falls back to the CLI, the fallback itself is now two-tiered:

  1. Tier 1 β€” warm pool (fast path): WorkerPool.complete(body, abortSignal, overflowFn) keys requests by model :: sha256(systemPrompt)[:16]. The first request for a given key lazily spawns a persistent claude -p --input-format stream-json --output-format stream-json subprocess (concurrency-1, FIFO queue) and reuses it warm on subsequent calls. ~2-3s steady-state.
  2. Tier 2 β€” overflow (cold one-shot): completeClaudeCodeViaCLI runs a cold one-shot execFile of claude -p per request (~10-14s, the legacy behaviour). The pool falls through to overflow when all workers for a key are busy, when the key is in crash-cooldown, or when the worker pool is disabled via LLM_PROXY_DISABLE_WORKER_POOL=1 (the GUARD-01 escape hatch).

Lifecycle

The pool guarantees four lifecycle behaviours (WLIFE-01..04):

Claude CLI Worker Pool lifecycle

  • WLIFE-01 β€” lazy spawn: zero workers exist at boot. The first claude-code fallback request spawns exactly one worker; idle keys cost nothing.
  • WLIFE-02 β€” idle eviction: a worker idle past LLM_PROXY_WORKER_IDLE_MS (default 30 min) disposes itself via an unref'd timer and is dropped, freeing RAM in quiet periods. The next same-key request lazily respawns it.
  • WLIFE-03 β€” crash recovery: a worker that crashes or EPIPEs mid-request surfaces its in-flight and queued jobs as RETRYABLE (never a hang). Per-key crash tracking puts a storming key into cooldown β€” routing it to overflow β€” once it crosses LLM_PROXY_WORKER_CRASH_THRESHOLD crashes within LLM_PROXY_WORKER_CRASH_WINDOW_MS, preventing a spawnβ†’crashβ†’respawn storm.
  • WLIFE-04 β€” cancellation: a client disconnect/abort SIGTERMs and disposes the in-flight worker and drops it synchronously, so the next same-key request gets a fresh cold worker. A queued (not-yet-running) abort only dequeues that one job, leaving the live worker untouched.

Workers also recycle proactively after LLM_PROXY_WORKER_MAX_REQUESTS requests or LLM_PROXY_WORKER_MAX_INPUT_TOKENS cumulative input tokens, and an LRU prompt-cap (LLM_PROXY_WORKER_PROMPT_CAP) bounds the number of distinct (model Γ— prompt) pools kept alive. See Worker Pool tuning for the full env-knob reference.


API Endpoints

GET /health

Returns proxy status and provider availability.

Response:

{
  "status": "ok",
  "providers": {
    "copilot": {
      "available": true,
      "mode": "direct-http",
      "lastChecked": 1708600000000
    },
    "claude-code": {
      "available": true,
      "version": "2.1.50",
      "lastChecked": 1708600000000
    }
  },
  "uptime": 3600,
  "inFlightRequests": 0
}

POST /api/complete

Forward a completion request to an LLM provider.

Request:

{
  "provider": "copilot",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain dependency injection."}
  ],
  "model": "claude-sonnet-4.5",
  "maxTokens": 1000,
  "temperature": 0.5,
  "tier": "standard"
}

Response (200):

{
  "content": "Dependency injection is...",
  "provider": "copilot",
  "model": "claude-sonnet-4.5",
  "tokens": {"input": 25, "output": 150, "total": 175},
  "latencyMs": 2100
}

Error Responses:

Status Type Meaning
400 VALIDATION_ERROR Missing required fields or unknown provider
401 AUTH_ERROR OAuth token expired or missing. For claude-code, the dispatcher will normally have already fallen back to the CLI before surfacing this.
429 QUOTA_EXHAUSTED Provider rate limit reached. For claude-code this surfaces only when both the direct OAuth bearer AND the CLI fallback return 429 β€” a true hard limit on the Max subscription. The common case (sonnet via bearer 429, CLI succeeds) is invisible to the caller.
503 PROVIDER_UNAVAILABLE Provider not configured or unreachable
504 TIMEOUT Request timed out
500 PROVIDER_ERROR Upstream error

Environment escape hatches (claude-code)

Env var Effect
LLM_PROXY_DISABLE_CLAUDE_DIRECT=1 Force every claude-code call through the legacy CLI path (skip the direct OAuth fast path). Useful if Anthropic tightens the bearer-endpoint allowlist or for debugging.

Configuration

Environment Variables

Variable Default Description
LLM_CLI_PROXY_PORT 12435 Port the proxy listens on
LLM_CLI_PROXY_URL - Set in Docker containers to http://host.docker.internal:12435

Docker Compose

The Docker container connects via host.docker.internal:

environment:
  - LLM_CLI_PROXY_URL=http://host.docker.internal:12435

Quick Start

# Start the proxy bridge (thin wrapper delegates to @rapid/llm-proxy)
node src/llm-proxy/llm-proxy.mjs

# Or use the standalone package directly
npx @rapid/llm-proxy

Auto-Start

The proxy starts automatically when launching coding (via bin/coding). The startup script (scripts/start-services-robust.js) handles:

  • Checking if already running on port 12435
  • Spawning the process with retry logic
  • Health check via GET /health

Health Monitoring

The proxy is monitored by the health verification system:

  • Severity: Warning (optional service)
  • Check type: HTTP health (GET /health)
  • Auto-heal: Disabled (host-side service, not Docker-managed)
  • Dashboard: Appears in System Health Dashboard service checks

Troubleshooting

Proxy not starting

  1. Check if port is already in use: lsof -ti:12435
  2. Check logs: node src/llm-proxy/llm-proxy.mjs

Copilot auth failure

OAuth tokens are read from ~/.local/share/opencode/auth.json. If expired:

# Re-authenticate via OpenCode
opencode  # tokens refresh on startup

Claude CLI not found

The CLI fallback path requires claude on the host:

which claude && claude --version

If the direct OAuth path works, the CLI is only needed for sonnet/opus calls (rate-limited on the bearer endpoint). Without the CLI, those calls will surface as QUOTA_EXHAUSTED.

claude-code calls take ~14s instead of ~1s

The dispatcher fell back to the CLI path. Check the proxy log:

tail -50 ~/Agentic/coding/.data/llm-proxy/logs/stdout.log | grep claude-code

Look for direct API rate-limited ... falling back to CLI. This is expected today for sonnet/opus β€” Anthropic's OAuth bearer endpoint per-model rate limit. The CLI fallback uses the same Max subscription but a different rate-limit bucket. As of v7.3 the Claude CLI Worker Pool keeps the fallback subprocess warm, bringing steady-state fallback latency down to ~2-3s; cold one-shot overflow calls still take ~10-14s.

"say OK" health-coordinator probe burns 16K tokens per call

This is the symptom of health-coordinator being routed through the CLI path. Verify the direct OAuth path is enabled:

env | grep LLM_PROXY_DISABLE_CLAUDE_DIRECT  # should be empty

If set to 1, unset it and restart the proxy:

launchctl kickstart -k gui/$(id -u)/com.coding.llm-cli-proxy

Healthy state: say OK calls report 9 input + 4 output tokens at ~0.9s.

Health dashboard shows degraded

Ensure the health check endpoint uses localhost:12435 (not host.docker.internal, which only resolves inside Docker).


Full Documentation

See the @rapid/llm-proxy package documentation for: