When parallel AI agents turn OAuth refresh-token rotation into an outage
Parallel agent workers can turn legitimate token refreshes into replay detections. Prevent the outage with per-grant coordination, atomic token-cache updates and provider-specific handling.
A hypothetical, but familiarIn a hypothetical agent service, two workers read the same shared refresh token as their cached access token expires. A provider using rotation with reuse detection accepts one refresh, then treats the other worker's stale token as replay and invalidates the replacement. Queued tasks stall when the service can no longer renew access.
- Confirm each provider's rotation and reuse-detection behaviour before sharing a refresh-token cache across workers.
- Serialise refreshes per shared grant across every worker and atomically publish the resulting token state.
- Treat Microsoft Entra's non-revocation of old tokens on use as documented provider behaviour, not a portable OAuth guarantee.
Reuse detection cannot distinguish theft from a coordination failure
RFC 9700, section 4.14.2, describes refresh-token rotation as a way to detect replay. The authorisation server issues a replacement refresh token, invalidates the previous token and retains their relationship. If the invalidated token appears again, the server cannot determine which party is legitimate and revokes the active refresh token. That response protects against a stolen token being reused, but it also means a legitimate client's coordination error can terminate its ability to renew access.
The race needs no malicious agent. Both workers read the same token before either publishes an update. The first refresh succeeds; the second reaches the provider after the previous token becomes invalid, outside any applicable grace period. Its stale request triggers reuse detection and invalidates the replacement already held by the first worker. AI is not a new OAuth mechanism here: parallel workers simply expose unsafe ownership of shared credential state.
Evidence: Best Current Practice for OAuth 2.0 Security RFC 9700 also known as BCP 240, Refresh access tokens and rotate refresh tokens
Entra replacement is not the same as reuse-triggered revocation
Microsoft's refresh-token documentation makes an important distinction: the Microsoft identity platform returns a fresh refresh token on use, but does not revoke the old refresh token when it is used to fetch new access tokens. Microsoft instructs clients to securely delete the old token after acquiring a new one. This is documented behaviour, not a recommendation to retain old tokens. Those tokens can still expire or be revoked for other reasons.
| Model | Previous refresh token | Documented consequence |
|---|---|---|
| RFC 9700 rotation | Invalidated | Replay revokes active refresh token |
| Okta rotation | Valid during configured grace | Reuse invalidates latest refresh token and access tokens issued since authentication |
| Microsoft Entra | Not revoked by refresh use | Replacement alone does not revoke it |
Okta documents a default 30-second grace period for rotation, configurable from zero to 60 seconds. Its stated purpose includes allowing recovery when newly issued tokens do not reach the client. My recommendation is to treat that allowance as a recovery margin, not a concurrency strategy. A delayed worker can outlive it, and a grace period does not make competing cache writes safe. Do not extrapolate Entra's behaviour to another provider.
Parallelise agent work, but serialise refreshes for each shared grant and make provider-specific token semantics an explicit design constraint.
Evidence: Best Current Practice for OAuth 2.0 Security RFC 9700 also known as BCP 240, Refresh tokens in the Microsoft identity platform, Refresh access tokens and rotate refresh tokens
Make one refresh owner responsible for each shared grant
My recommendation is a refresh lock scoped to the shared grant, enforced across every process or replica that can use it. A process-local mutex cannot coordinate separate workers. Use a stable grant or credential-record identifier, not the rotating token value, as the coordination key. Avoid resource-only partitioning: Entra documents refresh tokens as bound to user and client, not resource or tenant. Separate resource caches may therefore still depend on the same refresh credential.
- Identify every worker and replica that can read or refresh the shared credential.
- Use a cross-worker coordinator keyed to the stable shared-grant record.
- Acquire ownership before reading the refresh token used for the request.
- Re-read the cache after acquiring ownership and reuse a suitable access token if another worker has refreshed.
- Keep unrelated grants independent so refresh coordination does not serialise all agent work.
Evidence: Best Current Practice for OAuth 2.0 Security RFC 9700 also known as BCP 240, Refresh tokens in the Microsoft identity platform
The critical section ends after the cache commit
I would keep refresh ownership through the token-endpoint request and the durable cache update. Atomically publish the returned access token, any replacement refresh token, expiry metadata and cache version as one state change. Release ownership only after that commit succeeds. Otherwise, another worker can observe an incomplete update or refresh with the previous credential. This is an implementation recommendation derived from the rotation failure mode, not a locking requirement specified by RFC 9700.
I would also use version-checked writes to prevent an older worker from overwriting newer token state. That protects the cache, but it is not a substitute for refresh ownership: rejecting a stale write cannot undo a token request already accepted by the provider. Design lock-expiry and worker-crash handling with that distinction in mind. A replacement owner should not blindly send another refresh merely because the previous owner's lease has expired.
Evidence: Best Current Practice for OAuth 2.0 Security RFC 9700 also known as BCP 240, Refresh access tokens and rotate refresh tokens
Test overlapping refreshes and ambiguous failures
Make concurrency and failure handling explicit acceptance tests. For Okta, the guide documents the System Log events app.oauth2.as.token.detect_reuse for custom authorisation servers and app.oauth2.token.detect_reuse for the org authorisation server. These establish that reuse detection occurred; they do not establish whether the cause was credential theft or a client race. My recommendation is to correlate provider events with refresh ownership and cache-version records, without recording token values.
- Start competing workers against one grant and verify that only the owner submits a refresh.
- Delay a worker beyond any configured grace period and verify that it reloads shared state.
- Simulate response loss and a crash before cache commit to check recovery decisions.
- Attempt a stale cache write and confirm that it cannot replace newer token state.
- Verify that revoked or expired grants stop retrying and enter a defined recovery path.
Microsoft documents that refresh tokens can expire or be revoked and recommends graceful recovery through interactive sign-in. For an agent using a grant that requires that interaction, I would pause affected work and surface the reauthorisation requirement rather than run an unbounded retry loop. Entra's different replacement behaviour removes this particular reuse-triggered revocation assumption; it does not remove the need for coherent token state or an access-recovery plan.
Evidence: Refresh tokens in the Microsoft identity platform, Refresh access tokens and rotate refresh tokens
Sources and further reading
- Best Current Practice for OAuth 2.0 Security RFC 9700 also known as BCP 240
- Refresh tokens in the Microsoft identity platform
- Refresh access tokens and rotate refresh tokens
Source links support the documented product behaviour. Recommendations and labelled examples are editorial guidance.