OAuth integrations

OAuth refresh token races: trace the lost response

A rotating refresh token can be consumed even when its response is lost. Coordinate duplicate callers and guard local writes, then follow the provider's recovery contract for an uncertain exchange.

Model
Six constructed event schedules
Duplicate calls
Two exchanges revoke the strict modeled family
Write guard
Preserves a newly authorized grant
Limit
No HTTP or distributed coordination test
Two indigo credential cards approach an amber gate while one replacement card waits on the other side.
Conceptual illustration of competing refresh attempts. The article uses synthetic credentials in a local model.

The new token is stored, but the connection cannot refresh

Two workers read the same refresh token. The first exchanges it and stores its successor. The second sends the old value. Under a strict rotation policy with reuse detection, that second exchange can revoke the refresh-token family, including the successor already stored by the first worker.

This is a constructed event model, not a report of a customer incident. Its authorization server accepts each refresh token once and revokes the family when a consumed token returns. That policy lets us separate three failures that can produce similar connector symptoms: duplicate exchanges, an old response overwriting a new connection and an accepted exchange whose response never reaches the client.

The first trace is short:

  • Both callers read local version 0 and family-A:0.
  • Caller A sends family-A:0. The server consumes it and returns family-A:1.
  • A stores family-A:1 at local version 1.
  • Caller B sends the previously captured family-A:0.
  • The server detects reuse and revokes family A.

A successful database write at the third event does not prove that the credential remains usable after the fifth.

The stored value and the provider's current authorization state belong to different systems.

Both workers read family-A:0. Worker A exchanges it and stores family-A:1. Worker B sends the consumed token and revokes the modeled family.
Figure 1. Constructed strict-rotation sequence. Local persistence completes before the second caller causes remote revocation. View full-size figure.

Six schedules distinguish the failure boundaries

The downloadable Python event model executes six schedules and writes the complete traces. It uses synthetic token labels, explicit event order and assertions. There are no HTTP requests, real credentials or concurrent threads.

Six executed refresh schedules under strict one-use rotation
ScheduleToken-endpoint callsObserved result in the model
Duplicate callers2One replay event; family revoked; stored successor unusable
Coordinated callers1Waiter reads the committed successor; token usable
Late result, unconditional write1Old response overwrites a newly authorized grant
Late result, version guard1Old response rejected; newly authorized grant preserved
Lost response, blind retry2Retry triggers replay revocation
Lost response, stop1Family remains active remotely; client still lacks its successor

The paired results identify separate control points. Coordinating callers changes how many exchanges reach the provider. A version guard changes whether a response may replace the current local credential. Neither operation retrieves a successor that was issued remotely and then lost in transit.

In the late-response schedules, the user reconnects while an earlier refresh response is delayed. Reconnection stores family-B:0. An unconditional write then replaces it with the old grant's family-A:1. The guarded write compares the version captured before the exchange with the current version and rejects the mismatch. It preserves the new grant even though the old request completed successfully.

These are event-order demonstrations, not measured failure probabilities. The coordinated schedule assumes one effective owner. It does not test a distributed lock, its expiry or its behavior during a network partition.

Four selected OAuth schedules use two, one, two and one mock provider calls. Outcomes: revoked, usable, revoked and uncertain.
Four of six executed schedules. Fewer calls do not imply recovery: after a lost reply, stopping remains uncertain. Strict one-use policy, no real provider. View full-size figure.

Read the provider's contract before labeling a replay

RFC 6749, section 6 permits a refresh response to contain a new refresh token. When it does, the client must replace the previous value. A client also needs to handle a successful response that omits a replacement according to that provider's documented behavior. Replacing a stored refresh token with a missing value is a different implementation defect. For public clients, RFC 9700 requires replay detection through sender-constrained refresh tokens or rotation. Its rotation discussion describes invalidating the predecessor, retaining the relationship and revoking the active refresh token after reuse reveals a potential breach. The server cannot identify the legitimate party from that reuse alone.

An overlap window changes the concurrency outcome. Auth0's configuration documentation describes leeway during which the previous token may be reused without triggering breach detection. It also limits that exception to the previous token. Exchanging the second-to-last token triggers detection. The fixture deliberately has no overlap window. Its two-call revocation result therefore cannot be applied to every Auth0 configuration or every OAuth provider.

Record the actual connector contract: whether rotation occurs, which predecessor remains acceptable, how long any overlap lasts and what happens after reuse. Keep timeout recovery separate from those facts. A grace period is a provider feature with security consequences, not permission for an unbounded client retry loop.

Coordinate the connection, then guard the write

A single-flight operation lets several callers share one refresh attempt. Its key should identify the actual credential record, including the tenant and provider connection when those distinguish grants. A single global key can mix unrelated work. A key that is too narrow lets workers refresh the same grant independently.

The owner reads the credential again after acquiring coordination. A waiting worker may discover that another owner already installed a usable successor. It should consume that committed result rather than sending a refresh request from the snapshot it took before waiting.

Process-local coordination covers callers inside that process. Separate application instances, background workers and browser contexts need a design that accounts for every place able to exchange the shared credential. A distributed lease can reduce duplicate dispatches, but lease expiry can leave the earlier owner running while a new owner starts. Test that overlap explicitly. The background-job fencing guide explains the related stale-owner problem.

The write guard still matters after coordination. Store a version or connection epoch alongside the encrypted credential. Accept a refresh result only if the connection still matches the snapshot under which the request began. Reconnection and disconnection must advance the same guard so a delayed completion cannot resurrect an obsolete grant.

If a conditional write loses, read current state and determine why. A newer connection is different from another refresh winning the race. Do not respond to either by repeatedly sending the old refresh token. Local compare-and-swap protects stored state. It cannot undo a second token-endpoint call already accepted or rejected remotely.

A timeout leaves an unknown remote outcome

In the fifth schedule, the server accepts family-A:0, advances to family-A:1 and loses the response. The client still holds family-A:0. Retrying that value triggers the modeled server's reuse rule. Coordination did its job: there was only one original owner. The failure happened after remote acceptance.

The sixth schedule stops after the ambiguous timeout.

This avoids the additional replay event, but it does not restore the connection. The server's family remains active and the client has no usable refresh token. Marking the connection uncertain is an honest local state, not a recovery guarantee.

Recovery must follow a documented provider capability. If the provider supports a bounded retry of the predecessor, implement the exact scope and timing it specifies. If no supported recovery path exists, require a fresh authorization grant. Do not invent token-endpoint idempotency from the fact that other APIs accept idempotency keys. The API retry guide covers local duplicate suppression. The token endpoint remains a separate contract.

An invalid_grant response is also broader than a race diagnosis. Expiration, revocation and a wrong connection can produce rejected refreshes. Stop the immediate replay loop, then compare the request's connection version with provider evidence before attributing the error to concurrency.

Preserve evidence without recording the credentials

Useful diagnostic records include the connection ID, provider, local version read, dispatch attempt ID, completion outcome and whether the guarded write succeeded. Record transitions such as ready to refreshing or refreshing to uncertain. Keep access and refresh token values out of logs.

For the duplicate-call hypothesis, look for two dispatches associated with the same local credential version. For the stale-write hypothesis, look for a response completing after a newer authorization epoch was committed. For the lost-response hypothesis, identify an ambiguous transport failure and a subsequent attempt using an unchanged local version. Provider logs can help establish acceptance or reuse. A client timeout alone cannot.

The experiment tests refresh-token usability only. It makes no assertion about immediate access-token revocation at resource servers. That is another provider and resource-server policy to verify before promising that disconnection instantly stops every outstanding request.

Close the investigation with the uncertain path still visible

Run the fixture with Python 3 and inspect each trace alongside its assertions. The duplicate and blind-retry schedules must produce a replay event. The version guard must preserve the new grant. The stop-after-timeout case must remain uncertain with an unusable stored predecessor.

Then exercise the real integration using a controlled authorization server or an approved provider test environment. Delay the successful response until after reconnection. Trigger requests from separate workers. Terminate an owner after remote acceptance and before local persistence. Verify that waiting callers read current state, late results cannot replace the new connection and no generic retry middleware silently repeats the consumed token.

Keep the last test's expected outcome explicit: either the documented provider recovery succeeds, or the connection requests new authorization. An integration that preserves that distinction has evidence for its recovery behavior instead of treating every timeout as a failed exchange.

Sources

Documentation checked .

  1. IETF RFC 9700: OAuth 2.0 Security BCP, refresh-token protection
  2. IETF RFC 6749: refreshing an access token
  3. Auth0: refresh-token rotation and reuse detection
  4. Auth0: rotation overlap and its limits

Continue the conversation

Comments (1)

  1. Dreamtsoft Editorial

    Editorial follow-up: fewer refresh calls do not establish that a connection recovered. The lost-response scenario deliberately remains uncertain. Verify the provider’s documented recovery behavior before enabling automatic retries.

Leave a comment

Your name and comment stay in this page and are cleared after the spam check.

10–2,000 characters. Keep the discussion relevant to this article.

Spam protection verification
Spam protection loads when you begin the form.

JavaScript is required to use this form and its spam protection.