LLM Architecture¶
The provider layer underneath routing: which backends exist, how quota and fallback are handled, and what happens when a subscription runs out.
What this layer is¶
A single shared library that every LLM caller in the project goes through, replacing three separate abstractions that each handled providers differently. It owns provider registration, caching, circuit-breaking and metrics.
It answers how a provider is talked to. Which provider a given piece of work goes to is a separate question, answered by LLM Routing.
Subscriptions first, APIs as fallback¶
Work is served from flat-rate subscription accounts wherever possible, and only falls through to metered API keys when those cannot serve it — on quota exhaustion, or when a capability is missing.
Three modes¶
| Mode | Uses | For |
|---|---|---|
| Mock | No network at all | Tests and single-stepping a workflow |
| Local | A local model runner | Offline or cost-free work |
| Public | Cloud providers | Normal operation (default) |
Resilience defaults¶
A circuit breaker opens after 5 consecutive failures and resets after 60 seconds; an LRU cache holds 1,000 entries for an hour. Both are shared by every caller, so one service's retries do not become every service's outage.
The components¶
A facade is the single entry point for every LLM operation — it selects a provider, applies routing, and wires in the shared infrastructure. A registry holds the providers and validates their configuration. An infrastructure layer supplies the circuit breaker, the LRU cache and the metrics that the token-usage dashboard reads.

Three consumers sit on top: the semantic-analysis workflows, the inference engine, and the validator — all of which previously carried their own provider handling.
Subscription and metered accounts¶
The distinction that matters is not which company makes the model but which account gets billed. Subscription accounts are flat-rate: work served there costs nothing marginal. Metered API keys cost per token. The layer prefers the former and treats the latter as fallback, which is why quota tracking exists at all.
Beyond those, local runners and a mock provider round out the set — the mock being what makes a workflow single-steppable without spending anything.
Quota, backoff and falling through¶
Subscription quota is tracked optimistically: the layer assumes capacity until a provider says otherwise, rather than pre-counting. When a provider does refuse, exponential backoff keeps the system from hammering an exhausted account, and the call falls through to the next candidate.
The failure worth understanding is what happens when the only provider with a needed capability is exhausted — every chain that requires that capability collapses at once. The answer is to keep more than one capable provider configured, not to relax the capability check, because a provider that silently drops a capability returns confident nonsense.
Modes¶
Mock mode makes no network calls and is what the debug variant of the knowledge workflow uses. Local mode routes to a local runner. Public mode is the default and uses cloud providers.
Injection points¶
Three interfaces let you substitute behaviour without touching the layer: a mock service, a budget tracker, and a sensitivity classifier. They exist so that cost policy and data-handling policy can be project decisions rather than library ones.
What gets logged¶
Every call records its provider, model, calling process, token counts and latency. The process field is what makes per-service attribution possible on the Token Usage dashboard; a call without one is attributed to unknown.
Where to look next¶
- LLM Routing — which provider and model a given job resolves to
- Token Usage — what it all consumed
- LLM Providers — adding or troubleshooting a provider
Which layer this page describes
rapid-llm-proxy has two paths, and this page is about the lower one.
| Serves | Chooses a model by | |
|---|---|---|
| The library (this page) | Callers that embed the package directly — OKB and other consumers | fast / standard / premium tiers |
| The runtime bridge | Every LLM call in this project, on :12435 | a per-job route plus a small / medium / high band, from two YAML files |
Verified rather than assumed: the running daemon is proxy-bridge/server.mjs, which reads llm-routing.yaml and llm-fallback.yaml and imports only utility modules from the library's build — token accounting, cache parsing, network detection. It does not load the tier router described below.
So the tier model, the provider-priority chains and the quota logic on this page are real, and they are not what decides where your session's calls go. That is LLM Routing.
Overview¶
The @rapid/llm-proxy unified LLM layer consolidates three previously separate LLM abstractions into a single, comprehensive library that provides intelligent routing, resilience, and cost optimization across its provider implementations (12 concrete ones under src/providers/), including subscription-based providers that eliminate per-token API costs. It is maintained as a standalone package shared with OKB and other projects.
Why it was created: - Before: 3 separate LLM abstractions (Semantic Analysis, Unified Inference Engine, Semantic Validator) - Problem: Code duplication, inconsistent provider handling, no shared caching or resilience - After: Single unified layer (@rapid/llm-proxy) with shared infrastructure, tier-based routing, circuit breaker, and LRU cache
Key Features: - Parallelized copilot-first routing (library path) — Copilot scales well with parallelism (0.77s effective per call at 10 concurrent). The runtime bridge does not order providers globally: its defaults.background is claude-code-max and each route names its own provider - Zero-cost routing via GitHub Copilot and Claude Code subscriptions - Automatic fallback to paid APIs on quota exhaustion - Batch-optimized — agents already use Promise.all with concurrency 5-20, copilot as primary unlocks peak throughput - Optimistic quota tracking with exponential backoff
Architecture Components¶

Core Components¶
- LLMService (Facade)
- Single entry point for all LLM operations
- Handles tier-based routing
- Manages provider selection and fallback
-
Integrates with infrastructure (cache, circuit breaker, metrics)
-
ProviderRegistry
- Central registry for all LLM providers
- Dynamic provider registration and lookup
-
Configuration validation
-
Infrastructure Layer
- Circuit Breaker: Prevents cascading failures (threshold: 5 failures, reset: 60s)
- LRU Cache: 1000 entries, 1-hour TTL
- Metrics: Request tracking, cost monitoring, performance stats
Consumers¶
The unified layer serves three primary consumers:
- SemanticAnalyzer (
integrations/semantic-analysis/) - Batch analysis workflows
- Git history analysis
-
Ontology classification
-
UnifiedInferenceEngine (shared utility)
- General-purpose LLM inference
-
Multi-provider support
-
SemanticValidator (
integrations/constraint-monitor/) - Constraint violation detection
- Semantic code analysis
Supported Providers¶
The library ships 12 concrete provider implementations under src/providers/ with tier-based model selection:
Subscription Providers (Zero Cost)¶
1. Claude Code (Max subscription)¶
Path: Direct OAuth → CLI fallback (since 2026-05-19) Cost: $0 per token (uses existing Claude Max subscription)
| Tier | Model alias | Anthropic model returned | Typical latency (direct → fallback) |
|---|---|---|---|
| Fast | claude-haiku-4.5 | claude-haiku-4-5-20251001 | ~0.9s (direct path; rarely rate-limited) |
| Standard | claude-sonnet-4.6 | claude-sonnet-4-6-... | ~10-14s (typically falls back to CLI; bearer endpoint rate-limits sonnet) |
| Premium | claude-opus-4.6 | claude-opus-4-6-... | ~10-14s (same as sonnet — bearer rate-limited) |
Two-tier dispatch (proxy-bridge/server.mjs):
- Direct OAuth (fast path) —
completeClaudeCodeDirect()POSTs toapi.anthropic.com/v1/messageswithAuthorization: Bearer <oauth>read from the macOS keychain. Real token counts (9 in / 4 out forsay OK). Bearer +anthropic-version: 2023-06-01+anthropic-beta: oauth-2025-04-20. - CLI fallback (slower) — on 401/403/429, falls back to
completeClaudeCodeViaCLI()which spawnsclaude -p --output-format json --tools '' .... Uses the same Max subscription but a different Anthropic rate-limit bucket (sonnet/opus succeed via CLI when bearer 429s). Costs ~16-22Kcache_creationtokens per call because the CLI auto-injects its full system prompt + tool definitions.
Requirements: - The claude CLI on the host (only used for fallback path): which claude && claude --version - Authenticated: env -u ANTHROPIC_API_KEY claude auth status returns loggedIn: true, authMethod: "claude.ai", subscriptionType: "max" - Keychain entry Claude Code-credentials populated by the CLI
Token cache (proxy-side, in-memory): - Read on first claude-code call; cached until expiresAt - now < 60s - On 401: cache cleared, CLI fallback fires, the CLI rotates the keychain blob; next direct call reads the fresh token
Escape hatches: - LLM_PROXY_DISABLE_CLAUDE_DIRECT=1 — force every claude-code call through the legacy CLI path
Claude CLI Worker Pool (v7.3)¶
The v7.3 milestone restructures the CLI fallback described above into a two-tier dispatch backed by a warm worker pool (proxy-bridge/worker-pool.mjs, wired into the claude-code dispatcher in proxy-bridge/server.mjs). The legacy fallback paid the CLI's full system-prompt warm-up (~16-22K cache_creation tokens, ~10-14s) on every call because each call cold-spawned a fresh claude -p. The pool keeps that subprocess warm and reuses it, cutting steady-state fallback latency to ~2-3s. This affects claude-code only — copilot stays a direct HTTP POST and is never pooled.

WorkerPool / ClaudeWorker, keyed by (model × systemPrompt):
WorkerPoolowns a map of per-key worker pools, keyed bymodel :: sha256(systemPrompt)[:16]. Keying on both the model and the system prompt means callers sharing an identical (model, prompt) pairing reuse the same warm subprocess, while distinct prompts get isolated pools. An LRU prompt-cap bounds the number of distinct (model × prompt) pools held in memory.ClaudeWorkerwraps one persistentclaude -p --input-format stream-json --output-format stream-jsonsubprocess running concurrency-1 behind a FIFO queue. It owns that worker's lifecycle: lazy spawn, idle eviction, crash tracking, request/token-based recycling, and cancellation (see LLM Proxy Bridge → Claude CLI Worker Pool for the WLIFE-01..04 lifecycle guarantees).
How it slots under the claude-code dispatcher: the dispatcher's direct-OAuth fast path is unchanged. Only when it falls back to the CLI does the pool engage, as Tier 1: WorkerPool.complete(body, abortSignal, overflowFn). If every worker for a key is busy, the key is in crash-cooldown, or LLM_PROXY_DISABLE_WORKER_POOL=1 is set, the call transparently degrades to Tier 2 overflow — completeClaudeCodeViaCLI, the original cold one-shot execFile of claude -p. The pool is therefore a latency optimisation layered inside the existing CLI-fallback leg of the two-tier claude-code dispatch, not a new provider.
2. GitHub Copilot (Primary Provider)¶
Method: Direct HTTP POST to Copilot API Cost: $0 per token (uses existing GitHub Copilot subscription)
| Tier | Model | Description |
|---|---|---|
| Fast | claude-haiku-4.5 | Benchmarked: 5s sequential, 0.77s @10 parallel |
| Standard | claude-sonnet-4.5 | Claude Sonnet 4.5 via Copilot |
| Premium | claude-opus-5 | Claude Opus 5 via Copilot |
Why Copilot is primary: Performance benchmarks revealed that Copilot API calls scale beautifully with parallelism — 0.77s effective per call at 10 concurrent (vs 5s sequential). Since batch agents already parallelize LLM calls via Promise.all (concurrency 5-20), copilot as the first-choice provider unlocks peak throughput.
Authentication: - Reads OAuth token from ~/.local/share/opencode/auth.json - Direct HTTP POST to OpenAI-compatible Copilot API endpoint - No CLI tools required
Features: - Shared quota tracking system - Automatic provider rotation on exhaustion - Zero API costs - From containers: Falls back to LLM Proxy Bridge on host.docker.internal:12435
LLM Proxy Bridge (Docker Bridge)¶
When running inside Docker, host-side tools are unavailable. The LLM Proxy Bridge runs on the host (port 12435) and forwards requests. For Copilot, the proxy bridge reads OAuth tokens from ~/.local/share/opencode/auth.json and makes direct HTTP POST calls to the Copilot API. For Claude Code, it spawns the claude CLI. Each provider automatically detects and uses the proxy during initialization when the LLM_CLI_PROXY_URL environment variable is set.
API Providers (Per-Token Cost)¶
3. Groq¶
API Key: GROQ_API_KEY
| Tier | Model | Performance | Cost |
|---|---|---|---|
| Fast | llama-3.1-8b-instant | 750 tok/s | ~$0.05/M tokens |
| Standard | llama-3.3-70b-versatile | 275 tok/s | ~$0.59/M tokens |
| Premium | openai/gpt-oss-120b | - | High |
4. Anthropic¶
API Key: ANTHROPIC_API_KEY
| Tier | Model | Cost |
|---|---|---|
| Fast | claude-haiku-4-5 | $1/$5 per MTok |
| Standard | claude-sonnet-4-5 | $3/$15 per MTok |
| Premium | claude-opus-5 | $5/$25 per MTok |
5. OpenAI¶
API Key: OPENAI_API_KEY
| Tier | Model | Description |
|---|---|---|
| Fast | gpt-4.1-mini | Affordable small model |
| Standard | gpt-4.1 | Latest standard model |
| Premium | o4-mini | Reasoning model |
6. Google Gemini¶
API Key: GOOGLE_API_KEY
| Tier | Model | Description |
|---|---|---|
| Fast | gemini-2.5-flash | Fast, cost-effective |
| Standard | gemini-2.5-flash | Good balance |
| Premium | gemini-2.5-pro | Deep reasoning |
7. GitHub Models¶
API Key: GITHUB_TOKEN Base URL: https://models.github.ai/inference/v1
| Tier | Model |
|---|---|
| Fast | gpt-4.1-mini |
| Standard | gpt-4.1 |
| Premium | o4-mini |
8. DMR (Docker Model Runner)¶
Local provider - no API key required Base URL: http://localhost:12434/engines/v1
- Default Model:
ai/llama3.2 - Specialized Models:
ai/llama3.2:3B-Q4_K_M(lightweight tasks)ai/qwen2.5-coder:7B-Q4_K_M(code analysis)
9. Ollama¶
Local provider - no API key required
- Supports any locally installed Ollama models
- Used as final fallback for local-only mode
10. Mock Provider¶
Test/debug mode - no API key required
- Returns simulated responses
- Used for testing and development
Tier-Based Routing¶

Library path only
The tiers below are the library's routing model (ModelTier in src/types.ts, applied by src/provider-registry.ts). This project's own calls do not use it — they resolve through a per-job route and a small/medium/high band, with fallback given by an ordered per-provider chain rather than one global priority list. The fixed "Copilot → Groq → Claude Code → …" ordering stated below is therefore not the order your calls take; see LLM Routing for the one that is.
The library routes requests to providers based on task complexity and cost optimization:
The premium tier said claude-opus-4.6 until 2026-09-04
Every table below used to name claude-opus-4.6 (Copilot spelling) or claude-opus-4-6 (Anthropic wire spelling), because that is what the library declared — in src/config.ts, config/llm-providers.yaml, copilot-provider.ts and anthropic-provider.ts. The id was real and was genuinely served (3,268 token-usage rows via Copilot between 2026-03-06 and 2026-07-26), but it has since left both catalogues: opencode models now lists Copilot's opus ids as 4.7, 4.8 and 5, and llm-routing.yaml declares claude-opus-4.8 and claude-opus-5.
The library now names claude-opus-5 on both spellings — one string valid on the Anthropic wire and in Copilot's catalogue. The claude-code alias table further up still shows 4.6 on purpose: that column is the daemon's CLAUDE_OAUTH_MODEL_MAP, which still resolves the bare opus alias to claude-opus-4-6 and was not part of that change.
Tier Definitions¶
Fast Tier (zero cost → low cost, high speed) - Simple extraction and parsing - Basic classification - File pattern matching - Provider Priority: Copilot → Groq → Claude Code → Anthropic → OpenAI → Gemini → GitHub Models
Standard Tier (zero cost → balanced cost/quality) - Semantic code analysis - Git history analysis - Documentation linking - Ontology classification - Provider Priority: Copilot → Groq → Claude Code → Anthropic → OpenAI → Gemini → GitHub Models
Premium Tier (zero cost → highest quality) - Insight generation - Pattern recognition - Quality assurance review - Deep code analysis - Provider Priority: Copilot → Groq → Claude Code → Anthropic → OpenAI → Gemini → GitHub Models
Task-to-Tier Mapping¶
Tasks are automatically mapped to tiers based on their complexity:
# Fast tier examples
- git_file_extraction
- commit_message_parsing
- basic_classification
# Standard tier examples
- git_history_analysis
- semantic_code_analysis
- ontology_classification
# Premium tier examples
- insight_generation
- observation_generation
- pattern_recognition
Fallback Chain¶
- Primary: Try providers in priority order (copilot first — parallelism-optimized)
- Subscription check: Verify quota availability (copilot, claude-code)
- Circuit breaker check: Skip failed providers temporarily
- Cache check: Return cached results if available
- API fallback: Use paid API providers (Groq, Anthropic, OpenAI, Gemini, GitHub Models)
- Local fallback: DMR → Ollama (always available, no API costs)
Parallelism: Batch agents call LLMService via Promise.all (concurrency 5-20). Copilot scales from 5s sequential to 0.77s effective per call at 10 concurrent, making it ideal as the primary provider.
Subscription Quota Management¶
The system tracks subscription usage and automatically handles quota exhaustion:
Quota Tracking¶
Storage: .data/llm-subscription-usage.json
Tracked Metrics: - Completions per hour (rolling window) - Estimated token usage - Quota exhaustion state - Consecutive failure count
Soft Limits: - Claude Code: 100 completions/hour - Copilot: 100 completions/hour
Exponential Backoff¶
When quota is exhausted, the system applies exponential backoff:
- First exhaustion: Retry after 5 minutes
- Second exhaustion: Retry after 15 minutes
- Third+ exhaustion: Retry after 1 hour
Automatic recovery: On successful completion, reset failure counters
Automatic Fallback¶
Request → Check Copilot quota (primary — parallelism-optimized)
↓ (exhausted)
→ Use Groq (paid API, fast fallback)
↓ (circuit breaker open)
→ Check Claude Code quota
↓ (exhausted)
→ Use Anthropic (paid API)
↓ (all failed)
→ Use DMR (local)
Cost Impact: - If subscriptions available: $0 - If subscriptions exhausted: Standard API costs apply - Seamless transition - no user intervention needed
Data Persistence¶
Quota data is automatically: - Persisted to disk after each request - Pruned (keep last 24 hours only) - Loaded on service initialization
Reset quota tracking (for testing):
Mode Routing¶
The system supports three routing modes:
1. Mock Mode¶
Environment: SEMANTIC_ANALYSIS_MODE=mock
- Uses Mock provider exclusively
- Returns simulated responses
- No API calls or costs
- Ideal for testing and development
2. Local Mode¶
Environment: SEMANTIC_ANALYSIS_MODE=local
- Uses DMR and Ollama only
- No external API calls
- Zero API costs
- Requires local model servers running
3. Public Mode (Default)¶
Environment: SEMANTIC_ANALYSIS_MODE=public or unset
- Uses all cloud providers (Groq, Anthropic, OpenAI, Gemini, GitHub)
- Falls back to DMR/Ollama if all cloud providers fail
- Optimizes for quality and availability
Dependency Injection Hooks¶
The system provides interfaces for extending functionality:
1. MockServiceInterface¶
Used to customize mock responses for testing.2. BudgetTrackerInterface¶
interface BudgetTrackerInterface {
trackCost(provider: string, tokens: number, cost: number): void;
getBudgetRemaining(): number;
isOverBudget(): boolean;
}
3. SensitivityClassifierInterface¶
interface SensitivityClassifierInterface {
classifyContent(content: string): 'public' | 'internal' | 'confidential';
canUseProvider(provider: string, sensitivity: string): boolean;
}
Configuration¶
All LLM provider configuration is centralized in:
File: config/llm-providers.yaml
Key configuration sections:
providers:
# Subscription providers (zero cost)
claude-code:
cliCommand: "claude"
timeout: 60000
models:
fast: "sonnet"
standard: "sonnet"
premium: "opus"
quotaTracking:
enabled: true
softLimitPerHour: 100
copilot:
cliCommand: "copilot-cli"
timeout: 120000
models:
fast: "claude-haiku-4.5" # Benchmarked: 0.77s @10 parallel
standard: "claude-sonnet-4.5"
premium: "claude-opus-5"
quotaTracking:
enabled: true
softLimitPerHour: 100
# API providers (per-token cost)
groq:
apiKeyEnvVar: GROQ_API_KEY
fast: "llama-3.1-8b-instant"
standard: "llama-3.3-70b-versatile"
premium: "openai/gpt-oss-120b"
anthropic:
apiKeyEnvVar: ANTHROPIC_API_KEY
fast: "claude-haiku-4-5"
standard: "claude-sonnet-4-5"
premium: "claude-opus-5"
# ... more providers (openai, gemini, github-models)
# Copilot first — scales with parallelism (0.77s effective @10 concurrent)
# Batch agents use Promise.all, so copilot as primary unlocks peak throughput
provider_priority:
fast: ["copilot", "groq", "claude-code", "anthropic", "openai", "gemini", "github-models"]
standard: ["copilot", "groq", "claude-code", "anthropic", "openai", "gemini", "github-models"]
premium: ["copilot", "groq", "claude-code", "anthropic", "openai", "gemini", "github-models"]
cache:
maxSize: 1000
ttlMs: 3600000 # 1 hour
circuit_breaker:
threshold: 5
resetTimeoutMs: 60000 # 1 minute
Environment Overrides¶
Force specific behavior via environment variables:
# Force all tasks to premium tier
export SEMANTIC_ANALYSIS_TIER=premium
# Force specific provider (skip routing)
export SEMANTIC_ANALYSIS_PROVIDER=anthropic
# Use budget mode (fast tier everywhere)
export SEMANTIC_ANALYSIS_COST_MODE=budget
# Use local-only mode
export SEMANTIC_ANALYSIS_MODE=local
Cost Management¶
Cost Limits¶
Per-workflow limits (configured in llm-providers.yaml): - Budget mode: $0.05 per run - Standard mode: $0.50 per run - Quality mode: $2.00 per run
Batch workflow limits: - Max tokens per batch: 500,000 - Max cost per batch: $1.00 USD - Total budget: $50.00 USD - Automatic fallback to local on quota exceeded
Cost Optimization Strategies¶
- Copilot-first parallelized routing: Copilot scales with concurrency (0.77s @10 parallel), batch agents use Promise.all
- Tier-based routing: Use cheapest provider that meets quality requirements
- Caching: Avoid duplicate LLM calls (1-hour TTL)
- Automatic fallback: Switch to paid APIs only when subscriptions exhausted
- Local fallback: Switch to DMR/Ollama when budget exhausted
- Circuit breaker: Stop calling failed providers quickly
Cost Savings Example¶
Typical UKB batch analysis run: - 50 fast tier calls (extraction, parsing) - 100 standard tier calls (semantic analysis) - 20 premium tier calls (insight generation)
Before subscriptions (all API): - Fast: 50 × $0.001 = $0.05 - Standard: 100 × $0.01 = $1.00 - Premium: 20 × $0.05 = $1.00 - Total: ~$2.05 per run
After subscriptions (until quota exhausted): - Fast: $0 (Copilot/claude-haiku-4.5, parallelized) - Standard: $0 (Copilot/claude-sonnet-4.5) - Premium: $0 (Copilot/claude-opus-5) - Total: $0.00 per run ✅ - Bonus: ~3x faster via parallelized copilot calls
Savings: When the subscription quota covers the request, the per-token cost is $0 — material per-call savings vs. the paid-API fallback path.
Token Usage Monitoring¶
All LLM calls pass through the LLM Proxy Bridge (:12435), which logs every request to a persistent SQLite database for usage analytics and cost attribution.

How It Works¶
Every cognitive process that calls the LLM proxy includes a process identifier in the request body. The proxy logs the call with provider, model, token counts, latency, and subscription type to .observations/token-usage.db.
Logged Fields¶
| Field | Description | Example |
|---|---|---|
provider | LLM provider used | copilot, claude-code, anthropic |
model | Specific model | claude-sonnet-4-5, claude-haiku-4.5 |
process | Cognitive process identifier | observation-writer, consolidator, wave1-analysis |
input_tokens | Prompt tokens consumed | 1,200 |
output_tokens | Completion tokens generated | 450 |
latency_ms | Round-trip time | 2,340 |
subscription | Billing source | copilot-subscription, max-subscription, api-key |
Cognitive Process Identifiers¶
| Process | Description | Typical Volume |
|---|---|---|
observation-writer | Session observation classification + summary | High (every session event) |
consolidator | Digest → insight synthesis | Medium (batch, periodic) |
insight-generator | Entity insight generation | Low (batch, on refresh) |
content-validator | Entity content validation + refresh | Low (on demand) |
wave1-analysis | Code analysis agents | Medium (batch workflows) |
constraint-check | Constraint rule evaluation | Low |
auto-heal | Service recovery decisions | Rare |
health-check | Proxy liveness pings | High (every 30s) |
Dashboard¶
The Token Usage page on the Health Dashboard (:3032/token-usage) provides:
- Token distribution by process — treemap showing biggest consumers
- Provider breakdown — donut chart of provider usage
- Timeline — hourly token consumption area chart
- Recent calls — sortable table with process, model, tokens, latency
See Token Usage Dashboard for full documentation.
API Endpoints¶
The proxy exposes query endpoints for programmatic access:
GET /api/token-usage/summary?hours=24 # Aggregated stats
GET /api/token-usage/recent?limit=50 # Recent calls
Storage¶
- Database:
.observations/token-usage.db(SQLite) - Retention: All calls logged indefinitely
- Size: ~1KB per logged call
In-Memory Metrics¶
The LLM layer also tracks in-memory metrics per session:
- Request metrics: Total calls per provider, success/failure rates
- Performance: Latency per provider, throughput
- Cost: Token usage, estimated costs per provider
- Cache: Hit/miss ratio, cache size
Metrics are exposed via the LLMService.getMetrics() method.
Integration Example¶
import { LLMService, loadProviderConfig } from '@rapid/llm-proxy';
import type { LLMCompletionRequest } from '@rapid/llm-proxy';
const llmService = LLMService.getInstance();
const request: LLMCompletionRequest = {
messages: [
{ role: 'user', content: 'Analyze this code for bugs...' }
],
tier: 'standard', // or 'fast' / 'premium'
task: 'semantic_code_analysis',
temperature: 0.7,
maxTokens: 4096
};
const result = await llmService.complete(request);
Logger.log('info', result.content); // LLM response
Logger.log('info', `Provider: ${result.provider}`); // Which provider was used
Logger.log('info', `Model: ${result.model}`); // Specific model
Logger.log('info', `Tokens: ${result.usage.totalTokens}`); // Token usage
Logger.log('info', `Cached: ${result.cached}`); // Was it from cache?
Related Documentation¶
- Token Usage Dashboard - Detailed token usage monitoring documentation
- LLM Provider Guide - User guide for working with providers
- Semantic Analysis Integration - SA consumer usage
- Getting Started - Installation and API key setup
Configuration Files: - config/llm-providers.yaml - Full provider configuration schema - docs/provider-configuration.md - Detailed API key setup guide