AgentOnboard Docs

Production

Key caching and TTL, the last-known-good fallback and its cold-cache precondition, hourly refetch, rotation, and the operational caveats.

Everything here is a property of the key resolver in packages/sdk/src/keys.ts. The numbers are the ones in the source, not estimates.


What a request actually does

For every request, in order: read Authorization, decode the header, resolve a public key, verify, resolve the account. The only step that can touch the network is key resolution, and it is bounded:

token ──▶ decode header (kid)
             │
             ▼
        key in local cache for this kid?
             ├── yes ──▶ verify. No network.
             └── no  ──▶ is the cached set still fresh (≤ 1 hour old)?
                          ├── yes ──▶ serve from cache. No network.
                          └── no  ──▶ fetch the key set, shape-check it, cache it, verify
                                       └─ fetch failed ──▶ fall back to the last set that
                                                              passed the shape check

The SDK owns no socket. It constructs a URL and hands it to your runtime's fetch, so the only requirement it places on a host is that a fetch global exists. It imports no node:net, node:http, node:https, node:fs, and no framework.

Not 'zero network calls'

Verification is local, but it is not network-free. The cache has a 1-hour maximum age, so a long-lived process refetches its key set about once an hour, and a cold process fetches on its first request. Past the hour, a slow or dead key endpoint adds latency to requests rather than failing them outright — see the stall below.


Caching, and the last-known-good fallback

A fetched key set is shape-checked before anything is cached from it. A set that is not { keys: [...] } of usable entries — an RSA key without a modulus, an entry with no kid — is a failed fetch, never a bad token. The cached set is rebuilt from the validated entries, so nothing an endpoint sends beyond the keys does not survive.

When a fetch fails, the SDK verifies against the last set that passed validation rather than rejecting. This closes a gap the underlying library leaves open: that library serves its cached set only while it is fresh, so without this a one-minute blip on the key endpoint would reject every request even though a good set was still in memory.

The fallback has a precondition: a warm cache

There is nothing to fall back to on a cold start. KEY_SOURCE_UNAVAILABLE fires exactly when a fetch fails and no previously validated set is held — the first request of a fresh process, or the first request after the process restarts. If you scale to zero, or deploy often, that first request is the one that fails during an outage. Every other request in the same minute still verifies.

  • Maximum age: 1 hour (CACHE_MAX_AGE_MS). Past it, the next verification refetches.
  • Cooldown on an unmatched kid: 30 seconds (COOLDOWN_DURATION_MS). An unknown kid triggers at most one refetch per 30 seconds per process, which is what stops a stream of forged tokens from amplifying into a stream of key-endpoint requests.
  • At most 8 distinct key sources per process, least-recently-used evicted. Remote sources are keyed on the URL, so two issuers behind one fetchJwks never share a set. This bound is not configurable.

Rotation

Rotation is not something you configure; it is something you survive. A token minted under the old key keeps verifying as long as the key is still published, and a token under the new key fails with UNKNOWN_KEY until your process learns the new set.

The sequence, in practice:

  1. Tokens begin arriving with a new kid.
  2. Your process does not have it. It refetches the key set — the cache is either stale or the kid triggered the cooldown refetch.
  3. The new key is in the set. Verification succeeds. No request failed.

A kid that is no longer published is UNKNOWN_KEY immediately, including for a token that was valid a second ago. There is no grace period and no revocation list on the verification path. If you see UNKNOWN_KEY in traffic, the causes are, in order of likelihood: a rotation, a staging issuer configured without AON_ISSUER, or a forgery. The code does not distinguish them, because from inside the verifier it cannot.



What is fixed, and not configurable

You will meet these whether or not you configure them:

PropertyValue
Accepted algorithmRS256 only, pinned by the verifier. The token's own header never chooses it.
Required claimsexp, jti, sub, email.
Key selectionA kid in the header is required to select a key.
Key minimumRSA 2048 bits. A smaller key is refused in the resolver as KEY_SOURCE_UNAVAILABLE, not surfaced as a malformed token.
Clock skew5 seconds (clockToleranceSeconds, default 5).
Token lifetime5 minutes, minted by us, not negotiable per request.
Cached stateKey sets only, bounded at 8 sources. Verification results are never cached.
MAX_CACHED_KEY_SOURCES8, not configurable.

Deployment checklist

  • audience is a literal in your source, matching the hostname you serve on.
  • issuer is set — explicitly or via AON_ISSUER — on anything that is not production.
  • The request timeout on protected routes is longer than 5 seconds, or the process is warmed at startup.
  • ACCOUNT_REQUIRED answers 403 and points at your connect flow.
  • Your resolver returns null or undefined for an unknown email, and never creates or links one.
  • Logs carry code and jti, never the token or the Authorization header.
  • You alert on KEY_SOURCE_UNAVAILABLE and on a rise in UNKNOWN_KEY.
  • auth.md is published, and its code column matches what your responses actually return.

On this page