> ## Documentation Index
> Fetch the complete documentation index at: https://wundergraphinc-alberto-router-651-remove-netpoll.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Response Cache (Experimental)

> Cache fetches to avoid subsequent requests reaching upstream subgraphs.

<Danger>
  **Experimental feature.** Response caching is in alpha and is subject to change, and should not be used in production environments.
</Danger>

The router caches two kinds of subgraph response.

**Entity fetches.** An entity fetch is the request the router sends to a subgraph through the `_entities` root field to resolve fields of an entity that another subgraph owns. Entries are keyed per entity and per selection set, so asking for different fields of the same entity does not share an entry.

**Root query fetches.** A root query fetch is the request the router sends a subgraph for the root fields the client selected. The whole answer is stored as one entry. The key covers the operation text and every variable value. Any difference in either is a miss, and the request goes to the subgraph as it would have. Only queries are cached. Mutations and subscriptions are never cached.

Response caching is disabled by default.

## What gets cached

A subgraph response is cached only when its `Cache-Control` header asks to be. The rules apply in this order:

1. `no-store` refuses caching.
2. `no-cache` refuses caching, in any form. This includes the qualified `no-cache="Set-Cookie"` form.
3. `private` marks the response as belonging to a private context. It is cached only under a key scoped to that context, and only when `private_id` is configured and resolves for the request. See [Private responses](#private-responses). Without one, `private` refuses caching.
4. `s-maxage` sets the lifetime and takes precedence over `max-age`.
5. `max-age` sets the lifetime when `s-maxage` is absent.
6. When neither lifetime is present, `fallback_ttl` applies if the header contains `public`, `private`, `must-revalidate`, `proxy-revalidate`, `stale-if-error`, or `stale-while-revalidate`. The stale directives require a valid duration, such as `stale-if-error=60`.

`public` is optional. For example, `Cache-Control: max-age=60` and `Cache-Control: s-maxage=60` both allow caching for 60 seconds.

A value of `0` for the selected lifetime refuses caching. It does not fall through to the next rule. For example, `max-age=60, s-maxage=0` is not cached, while `max-age=0, s-maxage=60` is cached for 60 seconds. An invalid duration in a recognized directive also prevents caching.

A missing or empty `Cache-Control` header is never cached. Headers containing only unrecognized directives, such as `cdn-cache-control=60` or `immutable`, are also never cached. `fallback_ttl` does not apply to these responses.

Revalidation and stale-response directives opt into the fallback lifetime. The response cache does not revalidate entries or serve expired entries.

A response carrying GraphQL errors is never cached, whatever its `Cache-Control` header says.

<Info>
  Configuring the response cache does not make anything cacheable on its own. Subgraphs opt in with a recognized caching directive, such as `Cache-Control: max-age=60` or `Cache-Control: public`. If nothing appears to be cached, check the subgraph response headers first.
</Info>

## Request directives

A client controls the cache for its own request through the `Cache-Control` request header. Two directives are honored:

* `no-cache` is never answered from the cache. Every fetch of the operation goes to its subgraph, and the responses are stored as they are for any other request, so a `no-cache` request refreshes the entries it bypassed.
* `no-store` keeps the responses to this request out of the cache. Nothing is stored, but a warm entry still answers the request, as [RFC 9111](https://www.rfc-editor.org/rfc/rfc9111#section-5.2.1.5) allows.

Sent together, `no-cache, no-store` neither reads nor writes the cache. Every other request directive is ignored, including `max-age`, `min-fresh`, `max-stale` and `only-if-cached`, as is the `Pragma` header. A `Cache-Control` request header that does not parse is ignored as a whole.

<Warning>
  Any client can send these directives, and they cannot be turned off. A client that sends `no-cache` on every request bypasses the cache on every request, and its subgraphs see all of its traffic. A `no-cache` request is not [deduplicated](/router/request-deduplication) with identical requests in flight either.
</Warning>

## Enable it

### Redis

Every router replica shares one cache and entries outlive the process. Define a [storage provider](/router/storage-providers), then reference it by `provider_id`.

Redis 7.0 or newer is required to use this feature.

<CodeGroup>
  ```yaml config.yaml theme={"system"}
  storage_providers:
    redis:
      - id: 'my_redis'
        cluster_enabled: false
        urls:
          - 'redis://localhost:6379'

  response_cache:
    enabled: true
    key_prefix: 'cosmo_response_cache:'
    storage:
      provider: 'redis'
      provider_id: 'my_redis'
    all:
      fallback_ttl: 30s
  ```
</CodeGroup>

An unknown `provider_id` fails startup.

### Router memory

Each replica holds its own cache, nothing survives a restart, and no Redis is needed.

<CodeGroup>
  ```yaml config.yaml theme={"system"}
  response_cache:
    enabled: true
    storage:
      provider: 'memory'
      max_entries: 2048
    all:
      fallback_ttl: 30s
  ```
</CodeGroup>

`provider` defaults to `redis`, so caching in memory has to be asked for by name. This prevents a missing `provider_id` from quietly turning one shared cache into one cache per replica.

## Configuration reference

```yaml config.yaml theme={"system"}
response_cache:
  # Enable response caching. Default: false.
  enabled: true
  # Prepended to every cache key so entries cannot collide with anything else
  # sharing the Redis instance. Redis only, ignored by the memory provider.
  # Default: cosmo_response_cache:
  key_prefix: 'cosmo_response_cache:'
  # What the cache does for every subgraph without an entry under subgraphs.
  all:
    # Cache this subgraph's responses. Set to false to cache only the
    # subgraphs listed under subgraphs. Default: true.
    enabled: true
    # Lifetime for a response with a recognized caching directive but no lifetime.
    # A response naming its own max-age or s-maxage gets that instead.
    # Minimum 1s. Default: 30s.
    fallback_ttl: 30s
    # Expression yielding the id of the private context a request belongs
    # to, such as a user or a tenant, so private subgraph responses can be
    # cached per context. See Private responses.
    # Default: none, private responses are not cached.
    private_id: ''
  # Per-subgraph settings, keyed by subgraph name. An entry replaces all
  # entirely and is explicit: nothing is inherited and nothing is defaulted.
  # See Per-subgraph configuration.
  subgraphs:
    inventory:
      enabled: false
    products:
      enabled: true
      fallback_ttl: 5m
  storage:
    # redis or memory. Default: redis.
    provider: 'redis'
    # ID of a provider declared under storage_providers.redis.
    # Required unless provider is memory.
    provider_id: 'my_redis'
    # Memory provider only. Default: 10000.
    max_entries: 10000
  invalidation:
    # Which secondary indexes are built over cached entries, so an entry can
    # later be found by what it is about. Each defaults to true.
    # Index under the cache tags a subgraph declared.
    cache_tag: true
    # Index under the subgraph that answered.
    subgraph: true
    # Index entities under their __typename.
    type: true
    endpoint:
      # Serve the invalidation endpoint. Default: false.
      enabled: false
      # Its own address, not a route on the router's port. Default: localhost:5027.
      listen_addr: 'localhost:5027'
      # Default: /invalidation.
      path: '/invalidation'
      # Sent as the Authorization header, no Bearer prefix. At least 32
      # characters. Required when the endpoint is enabled.
      shared_key: ''
  cache_tag_header:
    # Return the cache tags of a response in a header a CDN can purge by.
    # See "CDN purging" below. Default: false.
    enabled: false
    # Header name. Cloudflare reads Cache-Tag, Fastly reads Surrogate-Key.
    # Default: Cache-Tag.
    name: 'Cache-Tag'
    # Separates tags in the value. Fastly expects a space. Default: ,
    delimiter: ','
    # Largest value emitted, in bytes. Tags that do not fit are dropped
    # finest first. The default is Cloudflare's limit. Default: 16384.
    max_bytes: 16384
```

`storage` is required when `enabled` is `true`.

Environment variables bypass the config schema validation. In each enabled scope, a `fallback_ttl` of zero or less fails startup, as does a `private_id` that does not compile.

## Per-subgraph configuration

`all` applies to every subgraph. An entry under `subgraphs`, keyed by subgraph name, replaces `all` for that subgraph:

```yaml config.yaml theme={"system"}
response_cache:
  enabled: true
  storage:
    provider: 'redis'
    provider_id: 'my_redis'
  all:
    fallback_ttl: 30s
    private_id: "request.auth.claims.sub"
  subgraphs:
    inventory:
      enabled: false
    products:
      enabled: true
      fallback_ttl: 5m
```

Here `inventory` is never cached, `products` keeps responses without a lifetime for five minutes, and every other subgraph keeps them for thirty seconds.

An entry is a full override and is explicit. Nothing under `all` reaches a subgraph that has an entry, and no key is defaulted:

* `enabled` must be set. An entry without it does not cache that subgraph. A `fallback_ttl` on its own does not enable anything.
* An enabled entry must set `fallback_ttl`, or startup fails naming the entry.
* An entry without `private_id` has none. So `products` above does not cache its `private` responses, although `all` names an id. Repeat the `private_id` in the entry to keep it.

Setting `all.enabled` to `false` turns the cache off for every subgraph without an entry, so only the entries that set `enabled: true` are cached. With no such entry the router starts and logs a warning, and nothing is cached.

An entry naming a subgraph the graph does not have is kept and logged at startup, since a feature flag graph may have it.

The invalidation indexes, the invalidation endpoint, and the cache tag header are not per subgraph.

## Private responses

A subgraph whose answer depends on who is asking marks it `Cache-Control: private`. A shared cache must not serve that answer to anyone else, but the router can still cache it for the context it was produced in, so the same context asking again is served without a subgraph request.

To do that the router needs an id for that context. `private_id` is an [expression](/router/configuration/template-expressions) evaluated once per request that yields it, such as a JWT claim or a request header. The id partitions the cache key: what it names is up to the expression, and it need not be a single user. A tenant, an organisation, or a role are all valid contexts, as long as every request that shares the id may share the response.

```yaml config.yaml theme={"system"}
response_cache:
  enabled: true
  storage:
    provider: 'redis'
    provider_id: 'my_redis'
  all:
    private_id: "request.auth.claims.sub"
```

Under `all` it applies to every subgraph. Under a `subgraphs` entry it applies to that subgraph alone, and an entry without one has none. See [Per-subgraph configuration](#per-subgraph-configuration).

With a `private_id` in place:

* A `private` response is stored under a key that includes the id, so each context has an entry of its own. A lookup with one id never finds an entry stored under another.
* A request for which the expression yields nothing, such as an anonymous request whose claim is `nil`, is not cached when the subgraph answers `private`. The subgraph is asked every time. This is not an error.
* `public` responses are unaffected. They are stored under the shared key and served to every request, whatever the id.
* With [`cache_control_policy`](/router/proxy-capabilities/adjusting-cache-control) or a `most_restrictive_cache_control` response header rule in place, a response served from a private entry reaches the client with `Cache-Control: private` and the entry's remaining lifetime, so a browser or CDN in front of the router does not share it either. Without one, the router sends no `Cache-Control` for the hit.
* Invalidation by subgraph, type, or cache tag drops private entries like any other.

The id is digested with SHA-256 before it becomes part of the key, so raw ids never reach the store. The expression must evaluate to a string. A value of any other type is reported through the response cache warning described under [Observability](#observability) and treated as no id.

<Warning>
  `private_id` partitions the cache. It does not by itself guarantee that a response is served only to those it was meant for: two requests with the same id share an entry, whatever else differs between them. Propagated headers are not part of the key unless the subgraph names them in `Vary`, so a `private` answer that also depends on a header the id does not account for is served to every request with that id. Name the header in `Vary`, fold it into the expression, or have the subgraph answer `no-store`.

  An expression that yields the same value for different contexts, or one that can be set by the client itself, lets one context be served another's entry. Prefer a verified claim over a plain header.

  A [`set` response header rule](/router/proxy-capabilities/subgraph-response-header-operations) on `Cache-Control` replaces what the subgraph sent, `private` included. A hit served from a private entry is treated the same as the origin response, so such a rule can hand the client a `public` policy for it.
</Warning>

## Varying responses

A subgraph whose answer depends on a request header, such as `Accept-Language`, names it in a `Vary` response header. The router then keeps one entry per set of values it sent, learned from the first response, the way an HTTP proxy cache does:

* The first request for an operation misses. The names in the response's `Vary` header are stored at the entry's key as a record, and the body is stored under a variant key derived from the values the router sent.
* A later lookup finds the record, digests the values this request sends the subgraph for those headers, and looks the variant up under the resulting key. A varying response costs two cache round trips per lookup. A response without `Vary` costs one, as before.
* Values are taken from the request as the subgraph receives it, after header propagation. A header named in `Vary` that the router does not forward counts as empty, so it never splits the cache.
* `Vary: *` is never cached, and neither is a response naming more than 100 headers.
* A header whose value differs per caller, such as `Cookie` or `Authorization`, yields one variant per caller. Nobody is served another caller's body, but the entries are rarely reused. A subgraph whose answer is per caller should say `private` instead.
* A subgraph that changes what it varies on adds the new set of names to the record on its next miss, newest first, rather than replacing it. A lookup tries one variant per set in that order and serves the first it finds, so variants stored under the old names stay reachable until they expire. A record keeps at most 8 sets; past that the oldest are dropped.
* With `engine.enable_multi_fetch` on, a merged entity fetch carries one `Vary` for the whole response, and the router applies it to every alias in it. An alias whose own answer would have varied on less is stored under the wider set, which splits its entries by headers they do not depend on. That costs extra misses, never a wrong body. A merged fetch follows records the same way a single fetch does, with one second round trip for all of its aliases together. The lever for the cost is on the subgraph side: emit the minimum `Vary` for what was actually resolved rather than a fixed union for the service.

## Limitations

These apply to the current alpha.

**A batch is served from the cache only when every entity in it is present.** One miss sends the whole batch to the subgraph, including the entities that were already cached. Entities are stored individually, so a later batch made up only of cached entities is a full hit.

**Forwarded request headers reach the cache key only through `Vary`.** The key covers the request the router renders for the subgraph, which is the operation text and its variables. Headers added by header propagation are applied after that and never reach the key on their own. Two requests that differ only by a propagated header share one entry unless the subgraph names that header in `Vary`, as described under [Varying responses](#varying-responses). A subgraph that varies by request header without saying so must use `no-store`, or `private` together with a `private_id` whose value accounts for that header. Omitting `public` does not prevent caching.

**Cache hits produce no subgraph telemetry.** A hit skips the subgraph fetch, so no subgraph span or metric is recorded for it. Subgraph request counts fall as the hit rate rises. Use them to confirm the cache is working, not to measure traffic.

**Merged entity fetches have their own rules.** With `engine.enable_multi_fetch` on, entity fetches to the same subgraph within one parallel wave are merged into a single request with aliased `_entities` fields. The response is taken apart again on the way into the cache, so each alias stores one entry per entity, the same as the unmerged fetch it replaces would have. From there:

* **Each alias is served all-or-nothing, like a batch.** A warm alias is left out of the next merged request. A cold one is fetched whole.
* **Only a request that finds every alias warm is a cache hit.** When some aliases are warm and others are not, the router still sends the request for the cold ones.
* **The reported `Cache-Control` is capped at the warm aliases.** The lifetime reported for a partly warm fetch is the life left on its warm aliases.
* **Declared cache tags are dropped.** `apolloEntityCacheTags` is one flat list with nothing to say which alias each entry belongs to. The entries are still indexed by subgraph and type, and invalidation by either still reaches them. Invalidation by cache tag misses them until they expire.
* **A failed request skips everything that depends on the merged fetch.** When the request for the cold aliases fails to reach the subgraph, the warm aliases keep their cached data and only the cold ones report the failure. Fetches that depend on the merged fetch are skipped as a whole, including those that depend only on a warm alias. Their fields stay `null` without an error of their own.

## Cache tags

A subgraph names what its response is about by returning a cache tag extension. Tags are recorded as a secondary index over the cached entries, so an entry can be found by what it is about rather than only by the key it happens to sit under.

The shape depends on what was fetched, because an entity fetch caches one entry per entity while a root query fetch caches its whole answer as one.

**Entity fetches** return `apolloEntityCacheTags`, an array of arrays holding one list per entity, in the same order as the `_entities` answered:

```json theme={"system"}
{
  "data": { "_entities": [{ "__typename": "User", "id": "1" }, { "__typename": "User", "id": "2" }] },
  "extensions": { "apolloEntityCacheTags": [["users", "user-1"], ["users", "user-2"]] }
}
```

**Root query fetches** return `apolloCacheTags`, one flat array:

```json theme={"system"}
{
  "data": { "homepage": { "title": "Hello" } },
  "extensions": { "apolloCacheTags": ["homepage", "featured"] }
}
```

An outer array whose length does not match the entities answered costs the whole response its tags,
rather than being zipped as far as the shorter of the two. Either extension is consumed by the router and
never forwarded to clients. The router does not record tags from a merged entity fetch. See [Limitations](#limitations).

## Invalidation

Cached entries expire when their TTL ends. Send a request to the invalidation endpoint on its separate port to remove entries earlier.

<Danger>
  **An entry written while an invalidation runs can survive it.** The index lists the entry before its value lands, so the request finds nothing to remove and leaves it in place. The entry then serves data fetched before the invalidation until its TTL ends. Send the request again once in-flight writes have settled. This is a known bug and a fix is in progress.
</Danger>

### Enable the endpoint

The endpoint only starts when the response cache itself is enabled.

```yaml config.yaml theme={"system"}
response_cache:
  enabled: true
  invalidation:
    endpoint:
      enabled: true
      shared_key: 'a-random-string-of-at-least-32-characters'
```

The endpoint speaks plain HTTP. Keep `listen_addr` on the loopback interface, or terminate TLS in front of it, so the shared key is not sent in the clear.

### Send a request

`POST` a JSON array to the endpoint, with the shared key as the `Authorization` header. There is no `Bearer` prefix.

```bash theme={"system"}
curl -X POST "http://localhost:5027/invalidation" \
  -H "Authorization: ${INVALIDATION_SHARED_KEY}" \
  -H "Content-Type: application/json" \
  -d '[
    { "kind": "subgraph", "subgraph": "products" },
    { "kind": "type", "subgraph": "products", "type": "Product" },
    { "kind": "cache_tag", "subgraphs": ["products", "reviews"], "cache_tag": "product-42" }
  ]'
```

Each element names one of three kinds.

| `kind` | Removes | Fields |
| - | - | - |
| `subgraph` | Every entry the subgraph answered | `subgraph` |
| `type` | Every entity of that `__typename` the subgraph answered | `subgraph`, `type` |
| `cache_tag` | Every entry the listed subgraphs tagged with it | `subgraphs`, `cache_tag` |

A successful request answers `202` with the number of cache keys removed:

```json theme={"system"}
{ "count": 12 }
```

The whole array is validated before anything is removed. One bad element refuses the request with `400` and removes nothing, so it can be corrected and resent as a whole. Each error names the element it is about:

```json theme={"system"}
{ "errors": [{ "index": 1, "kind": "type", "message": "a \"type\" request requires a type" }] }
```

### Restrictions

**Tags can be outdated.** With the Redis provider, an entry cached again under different tags stays listed under its earlier tags until it expires. Invalidating a tag the entry no longer carries still removes it. With the memory provider, an entry is listed under the tags of its latest write only.

**Everything is scoped to a subgraph.** A type or a cache tag is indexed per subgraph that answered, so every request names one. There is no way to invalidate a type or a tag across all subgraphs in one element. List the subgraphs instead.

**The memory provider is per replica.** Each router holds its own cache, so an invalidation reaches only the replica it was sent to. Do not use the in-memory cache with multiple router instances in production.

## CDN purging

A CDN in front of the router can cache whole responses and purge them by tag. Enable `cache_tag_header` and the router returns the tags of each response in a header the CDN reads. Set the header name and delimiter the CDN expects. Cloudflare purges by `Cache-Tag` with tags separated by a comma, the default. Fastly purges by `Surrogate-Key` with tags separated by a space.

```yaml config.yaml theme={"system"}
response_cache:
  enabled: true
  cache_tag_header:
    enabled: true
    name: 'Surrogate-Key'
    delimiter: ' '
```

<Note>
  The header names what a response is about. It does not tell the CDN to cache it. The CDN caches on the `Cache-Control` the router returns, which the [cache control policy](/router/proxy-capabilities/adjusting-cache-control) sets. On a hit, the lifetime reported for the served fetch is the time left on the entry, so the CDN never holds a response longer than the router does.
</Note>

### What the header carries

Every cacheable fetch of the request contributes three tiers of tag, coarsest first:

| Tag | One per | Purge removes |
| - | - | - |
| `subgraph-{name}` | Subgraph a cached fetch went to | Everything that subgraph answered |
| `type-{subgraph}-{__typename}` | Entity type an entity fetch returned | Every response holding an entity of that type |
| The tag as declared | Tag in `apolloCacheTags` or `apolloEntityCacheTags` | Every response the subgraph tagged with it |

A root query fetch has no single type and contributes no type tag. Declared tags are passed through unchanged. Tags are unique within the header.

The `subgraph-` and `type-` prefixes are not reserved. A declared tag such as `subgraph-products` is indistinguishable in the header from the generated one, and a purge by it removes both. Avoid those prefixes in declared tags so a header tag maps back to one invalidation element.

```
Cache-Tag: subgraph-employees,subgraph-mood,type-mood-Employee,moods,employee-1,employee-10
```

### When it is sent

The header is sent on cache hits and misses alike. A hit rebuilds it from what was stored with the entry, so both carry the same value. A fetch that was not cacheable contributes nothing, and a response with no cacheable fetch has no header.

The header is not sent on subscriptions, deferred or batched responses, or responses carrying subgraph errors. A response with subgraph errors is marked `no-store` instead.

The header does not depend on the `invalidation` indexes. It is built even when every index is off.

### Size

`max_bytes` caps the header value. When the tags do not fit, they are grouped coarsest first, subgraph then type then declared, and the finest that do not fit are left out. A purge by a coarse tag still reaches the response. The default of 16384 bytes is Cloudflare's limit on `Cache-Tag`.

A declared tag that contains the delimiter, or a byte not valid in a header, is left out.

`max_bytes` caps the whole value, not one tag. Fastly reads at most 1,024 bytes per key and ignores a longer key and every key after it, even when the value fits. Keep each declared tag at or below 1,024 bytes when purging through Fastly.

### Purge both caches

A purge at the CDN removes the response from the CDN only. The next request reaches the router, which may still serve the entry from its own cache. Send the same change to the [invalidation endpoint](#invalidation) as well. Each header tag maps to one `kind`:

| Header tag | Invalidation element |
| - | - |
| `subgraph-products` | `{ "kind": "subgraph", "subgraph": "products" }` |
| `type-products-Product` | `{ "kind": "type", "subgraph": "products", "type": "Product" }` |
| `product-42` | `{ "kind": "cache_tag", "subgraphs": ["products"], "cache_tag": "product-42" }` |

Invalidate the router first, then purge the CDN, so the CDN's refetch does not pick up the stale entry.

### Validation

Startup fails when `name` is not a valid header name or is `Content-Length`, `Content-Type` or `Transfer-Encoding`, when `delimiter` is empty or not valid in a header, or when `max_bytes` is not greater than zero.

## Observability

The router logs once at startup when the cache is enabled. With the Redis provider:

```
INFO  Response cache enabled  {"fallback_ttl": "30s", "subgraph_overrides": 0, "storage_provider": "redis", "key_prefix": "cosmo_response_cache:", "storage_provider_id": "my_redis"}
```

With the memory provider:

```
INFO  Response cache enabled  {"fallback_ttl": "30s", "subgraph_overrides": 0, "storage_provider": "memory", "max_entries": 2048}
```

`fallback_ttl` is the value under `all` and `subgraph_overrides` counts the entries under `subgraphs`.

Nothing is logged when the cache is disabled.

Cache failures never fail a request. When a read or a write fails, the router serves from the subgraph and logs a warning:

```
WARN  Response cache degraded, serving from the subgraph instead
```

That warning is sampled to one line per second. An unreachable cache produces one failure per cacheable fetch of every request in flight, and logging it unsampled would compound the outage.

## Related

* [Storage Providers](/router/storage-providers) defines the Redis instance the `redis` provider references.
* [Adjusting Cache Control](/router/proxy-capabilities/adjusting-cache-control) sets the `Cache-Control` the client and any CDN in front of the router receive.
* [Cache Warmer](/concepts/cache-warmer) is a different cache. It pre-populates the operation plan cache and does not cache subgraph responses.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.