Skip to content

Prompt Caching, End to End

Why cache-write is zero for some agents, why the cache-read band sits on a floor and then jumps β€” and why neither is a bug.

The one-sentence model

Caching never stores your text. It stores the provider's computation over a stable prefix of the prompt, so the next request beginning with the same prefix skips re-computing it. Everything else follows from that.

Why there is a prefix at all

An LLM API is stateless. Every turn re-sends the whole conversation β€” system instructions, tool definitions, all prior messages, then the new input. The front of that barely changes between turns; history only grows at the end. That unchanging front is the only thing that can be cached.

A prefix runs from the very start. The first token that differs ends the reusable span β€” so changing something in the middle makes everything after it fresh again.

Why agents differ

Your prompt reaches a provider over one of two API formats, and the two have different caching capabilities. Which one is used depends on the provider your work routed to, not on anything you chose.

That is the whole explanation for cache-write being 0 on one agent and not another.

When reading totals

total_tokens counts fresh input plus output. Cache reads and writes are additive, in their own columns. For a cache-heavy foreground session the raw total can understate real consumption by more than two orders of magnitude β€” always add the cache columns back.

Every turn re-sends everything

The provider keeps nothing between requests, so each turn transmits the entire conversation so far as one prompt. The system instructions and tool definitions are identical every time, and the history only ever grows at the end, so the front of the prompt is stable and the tail is fresh.

Caching operates only on prefixes, because the provider hashes from the beginning: the first differing token ends the reusable span.

This is why an edit near the start of a conversation is expensive in a way that an edit at the end is not β€” it invalidates the cached computation for everything after it.

The wire decides what is possible

Requests go through the proxy on :12435, which speaks to providers in one of two API formats. The format β€” the wire β€” is what determines the caching behaviour you observe:

Anthropic wire OpenAI wire
Cache writes Explicit, and reported Not surfaced the same way
How fresh input is counted Excludes cache reads Prompt tokens include them

That second row is the one that bites. The same conceptual number is reported differently by the two wires, so a column holding both conventions cannot be summed across providers without normalising first β€” which the proxy now does at the parse boundary, subtracting cached tokens from the OpenAI-side prompt count so that cache columns are additive everywhere.

Reading the bands

A cache-read band sitting on a flat floor and then jumping, with no write in between, is the expected shape: the floor is the stable prefix being re-read every turn, and the jump is that prefix growing when enough new conversation has accumulated to be worth caching.

An agent showing zero cache writes is not failing to cache. It is on a path where writes are not surfaced, while reads still are.

What this means for cost

Two practical consequences.

Displaying consumption requires adding the cache columns back. total_tokens answers "what did we newly send and receive", not "what did this cost". A cache-heavy interactive session can read as a small fraction of its real consumption otherwise β€” the measured gap on one 24-hour window was large enough to invert which of two consumers looked dominant.

Stability at the front of the prompt is worth money. Anything that reorders or rewrites system instructions or tool definitions between turns discards the cached prefix and pays full price for it again.

Why this page exists

Open the dashboard's Performance β†’ context modal and two things look odd:

  • Cache write is 0 (or N/A) for some agents but not others.
  • The green "cache read" band mostly sits on a flat floor, then sometimes jumps higher β€” without any write ever happening.

Neither is a bug. Both fall out of where each agent's prompt travels and which caching rules apply on that path. This page explains the whole chain β€” for Claude Code, opencode, and GitHub Copilot β€” so that after reading it there are no questions left.

The one-sentence model

Caching never stores your text β€” it stores the provider's computation over a stable prefix of the prompt, so the next request that begins with the same prefix skips re-computing it. Everything below is a consequence of that single idea.


1. The prompt is rebuilt and re-sent every turn

An LLM API is stateless. The provider keeps nothing between requests. So on every single turn, the agent re-sends the entire conversation so far β€” system instructions, tool definitions, all prior messages, and the new user input β€” as one big prompt.

flowchart LR
    subgraph T1["Turn 1"]
        A1["System + Tools<br/>(stable)"] --> B1["History<br/>(short)"] --> C1["New input"]
    end
    subgraph T2["Turn 2"]
        A2["System + Tools<br/>(stable)"] --> B2["History<br/>(grew)"] --> C2["New input"]
    end
    subgraph T3["Turn 3"]
        A3["System + Tools<br/>(stable)"] --> B3["History<br/>(grew more)"] --> C3["New input"]
    end
    T1 --> T2 --> T3

The front of the prompt barely changes turn to turn (same system, same tools, and history that only ever grows at the end). That unchanging front is the stable prefix β€” the only thing that can be cached. The new bit at the end is fresh every time.

Definition: prefix

A prefix is a run of tokens from the very start of the prompt. Caching only works on prefixes because the provider hashes from the beginning: the first token that differs from a cached entry ends the reusable span. Change something in the middle and everything after it is "fresh" again.


2. The "wire": which API language the provider is spoken to in

Your prompt does not go straight to a model. It goes through the rapid-llm-proxy (localhost:12435), which forwards it to a provider. The provider is spoken to in one of two API formats β€” we call the chosen format the wire, because it's what actually goes over the network.

flowchart TD
    Agent["Agent<br/>(Claude Code / opencode / copilot)"] --> Proxy["rapid-llm-proxy :12435"]
    Proxy -->|"Anthropic wire<br/>POST /v1/messages"| Anth["Anthropic API<br/>(claude-code OAuth or anthropic key)"]
    Proxy -->|"OpenAI wire<br/>POST /chat/completions"| Cop["GitHub Copilot gateway<br/>(OpenAI-compatible)"]

The wire matters because the two formats have different caching capabilities:

Anthropic wire (/v1/messages) OpenAI wire (/chat/completions)
Who serves it here Anthropic API (Claude direct / Max OAuth) GitHub Copilot gateway
How caching is triggered Explicit β€” you place cache_control markers in the request Implicit β€” the gateway auto-detects a repeated prefix; nothing to set
Cache read reported? βœ… cache_read_input_tokens βœ… cached_tokens
Cache write reported? βœ… cache_creation_input_tokens ❌ no such field exists in the schema
Pricing signal read β‰ˆ 0.1Γ—, write β‰ˆ 1.25Γ— of input read is cheaper; write is invisible

The single most important row is the last-but-one: the OpenAI wire has no field for "cache write." So on that wire, cache-write can only ever be reported as 0 / N/A β€” not because nothing was cached, but because the protocol has no way to say it wrote a cache. This is the whole reason the number differs between agents.


3. cache_control breakpoints, explained plainly

On the Anthropic wire, caching is opt-in: you must mark where the stable prefix ends. That marker is a small tag attached to a piece of the prompt:

{ "type": "text", "text": "...the system prompt...",
  "cache_control": { "type": "ephemeral" } }

Think of it as a bookmark that says "store everything up to here." We call each bookmark a breakpoint. The rules are simple:

  • Read happens automatically when a new request's prefix matches a stored one (within the TTL β€” ~5 min, up to 1 h).
  • Write happens when you place a breakpoint at a new boundary that wasn't stored yet β€” the provider computes that span once and stores it, billing it at ~1.25Γ—.
  • Up to 4 breakpoints; a span below the minimum size (~1024 tokens) is silently ignored.
flowchart LR
    S["System + Tools"] --> H["History so far"]:::bp --> N["New input"]
    classDef bp stroke:#e0a300,stroke-width:3px
    H -. "cache_control breakpoint<br/>= store up to here" .-> H

Turn over turn, this produces the write-then-read rhythm that makes long sessions cheap:

sequenceDiagram
    participant A as Agent
    participant P as Provider (Anthropic wire)
    A->>P: Turn 1 β€” prefix marked with breakpoint
    P-->>A: cache WRITE (store prefix, ~1.25Γ—)
    A->>P: Turn 2 β€” same prefix + a bit more
    P-->>A: cache READ of turn-1 prefix (~0.1Γ—)<br/>+ WRITE of the newly-added span
    A->>P: Turn 3 β€” same again + a bit more
    P-->>A: cache READ of the bigger prefix<br/>+ WRITE of the new span

Each turn reads the prefix it already stored and writes only the small new slice β€” so a long conversation is dominated by cheap reads.


4. Under the hood: how cached and new content mix (and why it's exact)

The last section said caching stores "the provider's computation." This section makes that concrete β€” because understanding it removes the last bit of mystery: how can old content be reused and still combine correctly with new content?

What the model actually computes

A transformer reads the prompt one token at a time. For every token, at every layer, it computes two vectors β€” a Key and a Value (K/V). Intuitively:

  • the Key says "here is what I'm about, so later tokens can find me,"
  • the Value says "here is the information I'll hand over when someone attends to me."

Producing the answer works by attention: each new token looks back over the K/V of all earlier tokens, decides which are relevant, and pulls in their Values. Computing K/V for a long prompt is the expensive part β€” it grows with the number of tokens. That K/V computation is exactly what caching stores and skips.

So 'the cache' is not text β€” it's the K/V tensors

A cache hit means: load the prefix's already-computed K/V from the provider's memory instead of recomputing them. The bytes of your prompt still cross the wire every turn; what's saved is the re-computation, which is why cache-read is ~0.1Γ— the price, not free.

How the new content is mixed in

On a cache hit the provider loads the prefix K/V, computes K/V only for the new tokens, appends them, and runs attention over the whole sequence:

flowchart LR
    subgraph Cached["Stable prefix β€” cache HIT"]
        kv1["K/V for<br/>system + tools"]
        kv2["K/V for<br/>older history"]
    end
    subgraph Now["New this turn β€” computed now"]
        kvn["K/V for<br/>new input"]
    end
    kv1 --> Cat["one continuous K/V sequence"]
    kv2 --> Cat
    kvn --> Cat
    Cat --> Attn["Attention:<br/>new tokens attend over ALL K/V<br/>(cached + fresh, indistinguishably)"]
    Attn --> Out["next token"]

The model sees one seamless sequence. It cannot tell which K/V were loaded from cache and which were just computed β€” they are the same numbers either way. The generated output is bit-for-bit identical to running with no cache at all.

Why reuse is exact, not an approximation

This is the key insight, and it's the reason caching is safe to do at all. Transformers use causal (masked) attention: a token may only attend to tokens at or before its own position β€” never to anything that comes later.

flowchart LR
    t1["tok₁"]:::pfx --> t2["tokβ‚‚"]:::pfx --> t3["tok₃"]:::pfx --> t4["tokβ‚„<br/>(new)"]:::new
    t4 -. "attends back" .-> t3
    t4 -. "attends back" .-> t2
    t4 -. "attends back" .-> t1
    classDef pfx fill:#1f6f3f,color:#fff
    classDef new fill:#1f4f8f,color:#fff

Because attention only ever points backward, a prefix token's K/V depend only on the tokens before it β€” never on what gets appended later. Add a thousand new tokens at the end and the prefix's K/V are unchanged, down to the last digit. That independence is what makes reuse exact:

  • Why a cached prefix is always valid β€” appending new content can't alter earlier K/V, so the stored values stay correct.
  • Why caching only works on a prefix β€” change one token in the middle and every token after it now attends to something different, so all their K/V change and nothing past that point can be reused. (This is the same reason section 1 said the first differing token ends the reusable span.)

One line to remember

Cache = the prefix's K/V attention state. Causal masking makes those values independent of anything appended later β€” so they can be stored once and reused exactly, while new tokens are computed fresh and slotted in right after them.


5. The three agents, side by side

Same idea, three different paths. This is where the dashboard differences come from.

flowchart TD
    subgraph CC["Claude Code"]
        cc1["Sets cache_control itself"] --> cc2["Anthropic wire /v1/messages"] --> cc3["Anthropic API"]
    end
    subgraph OC["opencode (current default)"]
        oc1["OpenAI client, sets nothing"] --> oc2["OpenAI wire /chat/completions"] --> oc3["Copilot gateway"]
    end
    subgraph GC["GitHub Copilot"]
        gc1["OpenAI client, sets nothing"] --> gc2["OpenAI wire /chat/completions"] --> gc3["Copilot gateway"]
    end
Claude Code opencode (default routing) GitHub Copilot
Wire Anthropic /v1/messages OpenAI /chat/completions OpenAI /chat/completions
Provider Anthropic (Max OAuth / key) Copilot gateway Copilot gateway
Caching kind Explicit cache_control Implicit (gateway) Implicit (gateway)
Who sets breakpoints Claude Code sets its own nobody β€” not supported on this wire nobody β€” not supported on this wire
Cache read shown βœ… real cache_read βœ… cached_tokens (green) βœ… cached_tokens (green)
Cache write shown βœ… real cache_write (amber) ❌ N/A (schema has no field) ❌ N/A (schema has no field)

So the write column is honest, not broken

  • Claude Code writes to cache and the Anthropic wire reports it β†’ you see amber.
  • opencode / Copilot do benefit from caching (the Copilot gateway caches prefixes implicitly, which is the green you see) β€” but the OpenAI wire has no field to report a write, so the dashboard correctly shows N/A, never a misleading 0.

6. Reading the "Per-turn tokens" chart

Here is the actual modal (Performance β†’ context β†’ Explain on any run) that prompted this whole page:

Per-turn tokens chart from the context modal β€” a green cache-read floor at ~40k with fresh-input blue stacked above across ~900 turns, and summary cards reading cache read 44.4M, cache write 0, fresh input 44.2M, output 510k

The green cache read band sits on a flat ~40k floor (system + tools, reused every turn) and occasionally climbs higher where recent history is still warm; blue fresh input stacks on top. The cards below total it up: 44.4M cache read reused, 44.2M fresh input, and a cache write of 0 β€” the exact pattern this page explains (this run rode a wire that reports no writes, so reuse shows up only as green).

Every bar is the tokens transmitted that turn, coloured by how the provider billed them:

Colour Meaning Appears on
🟒 green cache read β€” prefix reused (~0.1Γ—) both wires
🟠 amber cache write β€” prefix stored this turn (~1.25Γ—) Anthropic wire only
πŸ”΅ blue fresh input β€” never cached, full price both wires
🟣 purple output β€” the model's generated tokens both wires

Why the green floor sits at ~40k, then sometimes jumps

The green floor is the part of the prefix that stayed byte-identical long enough to still be warm in the cache:

flowchart LR
    subgraph Warm["Reused as cache read (green)"]
        sys["System + Tools<br/>(~40k, never changes)"] --- hist["Older history<br/>(now stable, still within TTL)"]
    end
    subgraph Fresh["Billed fresh (blue)"]
        new["Recent turns + new input"]
    end
    Warm --> Fresh
  • The ~40k floor = system instructions + tool definitions. These are identical every turn, so they're always a cache hit.
  • Green jumps above the floor when conversation history that was fresh a few turns ago has now stopped changing and is replayed within the TTL β€” so it, too, matches the warm entry and the green boundary creeps up.
  • Green drops back to ~40k after a pause longer than the TTL (the cache expired) or after something invalidated the middle of the prefix β€” only the static head survives; everything after it is re-billed as blue.

That is the complete answer to "why is there green (reuse) even though write is 0?" β€” on the OpenAI wire the reuse is the gateway's implicit caching, reported as read; the write simply has no field to appear in.


7. What the proxy now does (and doesn't)

The rapid-llm-proxy was updated so that any traffic it forwards on the Anthropic wire gets cache_control breakpoints injected automatically β€” even when the agent didn't set them.

  • It injects two breakpoints: at the end of the system prompt and on the last message block (the stable-prefix boundary that grows each turn).
  • It's pure and safe: below-minimum spans are ignored by the API; disable with LLM_PROXY_DISABLE_PROMPT_CACHE=1.
  • Both native completion paths then read cache_read_input_tokens / cache_creation_input_tokens back into the token record.
flowchart LR
    In["Request forwarded on<br/>Anthropic wire (no breakpoints)"] --> Inj["Proxy injects cache_control<br/>at system end + last message"] --> Out["Provider writes + reads cache"]

What this does NOT change

This only affects traffic on the Anthropic wire. opencode and Copilot route on the OpenAI wire to the Copilot gateway, which the proxy cannot add cache_control to (the field doesn't exist there). Their N/A write stays N/A. To make an opencode session show real cache-write, route it to an Anthropic-wire provider (a rapid-proxy/<claude-model> model instead of github-copilot/<model>) β€” a configuration choice with its own cost/latency/rate-limit trade-offs.


8. Quick answers

Q: Why is cache write 0 for opencode but not Claude Code? opencode rides the OpenAI wire (Copilot gateway), whose usage schema has no cache-creation field. Claude Code rides the Anthropic wire, which reports writes. Different wire β†’ different reportable numbers.

Q: If nothing was written, why is there green (cache read)? The Copilot gateway caches repeated prefixes implicitly and reports them as cached_tokens. Reuse still happens; only the write side is unreportable on that wire.

Q: Why does the green sometimes exceed the ~40k floor with no writes? Recent history became stable and is still within the cache TTL, so a longer prefix matches the warm entry. TTL expiry drops it back to the static head.

Q: Is N/A the same as 0? No. 0 would mean "measured, and it was zero." N/A means "this wire cannot report it." The dashboard shows N/A on the OpenAI wire deliberately.

Q: How do I get real cache economics for opencode? Point it at a rapid-proxy/<claude-model> model so it uses the Anthropic wire; the proxy then injects breakpoints and the provider reports both read and write.


Everything on this page is visible per-run in the dashboard's Context & caching explainer β€” the billed-by-cache breakdown, the explicit-vs-implicit wire explanation, and the per-turn stacked bar chart:

Caching detail in the Context & caching explainer β€” billed-by-cache bar, wire explanation, per-turn chart, wire table

See also