Repository navigation
fix(client): back off when Retry-After is zero or not in the future - #258
Conversation
The retry interceptor honored Retry-After literally, so `Retry-After: 0` produced a 0 ms delay. Confluence Data Center's rate limiter sends `Retry-After: 0` on 429 responses (and on 200s), so every retry fired immediately, hit the empty token bucket again, and all attempts were used up within milliseconds. Observed on Data Center 9.2.9 (3 requests per second): `children --recursive` sent 36 requests in 186 ms and failed with "Rate limit exceeded". retryDelayMs now honors Retry-After only when it names a positive wait: a positive number of seconds or an HTTP-date in the future, still capped at 60 s. A missing, unparsable, zero, negative or already-past value is treated like a missing header and falls back to the existing exponential backoff (1 s, 2 s, 4 s with the default base delay), which is itself capped at 60 s. The x-ratelimit-* headers are deliberately not used as a hint: they add nothing while Retry-After is positive, and the 1 s base backoff already spans the 1 s refill interval reported alongside `Retry-After: 0`. Tests cover zero, numeric zero, whitespace, negative and unparsable values, past and current HTTP-dates, positive seconds and future HTTP-dates (still honored and capped), the backoff cap, and the interceptor wiring (sleep durations for a 429 with `Retry-After: 0` and with `Retry-After: 3`) without waiting in real time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
pchuri
left a comment
There was a problem hiding this comment.
Thanks for the clear diagnosis and focused fix. Falling back to the existing exponential backoff when Retry-After provides no usable positive delay makes sense, while preserving positive delays and the existing cap.
I reviewed the current head and ran the client test suite (268/268 passing) and lint successfully. I did not find a blocking issue in the reviewed changes. I also appreciate the explicit note that recursive traversal can still exceed a low rate limit: concurrency and retry-budget changes can remain a separate follow-up rather than expanding this PR.
I have not independently repeated the live-server/simulation checks or run the full repository test suite yet. This looks ready to move toward merge after final validation. Thanks again!
## [2.25.9](v2.25.8...v2.25.9) (2026-10-01) ### Bug Fixes * **attachments:** make --replace work on Data Center/Server ([#257](#257)) ([e9163e4](e9163e4)) * **client:** back off when Retry-After is zero or not in the future ([#258](#258)) ([5d9b3b9](5d9b3b9)) * **deps:** update axios to 1.20.0 to unblock security audit ([#263](#263)) ([1f62546](1f62546))
|
🎉 This PR is included in version 2.25.9 🎉 The release is available on: Your semantic-release bot 📦🚀 |
…#272) Each request backed off on its own, so a burst of concurrent requests (page tree traversal runs ten at a time) that hit a low shared limit, such as 3 requests per second on Data Center, was rejected together, retried in lockstep, and exhausted its retries even though a patient client would have finished. #258 fixed immediate retries on Retry-After: 0 but not this. Add a RequestGate shared by every request of a client. It handles 429 only; a 503 keeps its independent per-request retries. A new 429 doubles the minimum spacing between request starts (100 ms, capped at 2 s) and each success shrinks it by 10%. Only a positive Retry-After pauses all requests. A rejection counts against a request own retries only when it is new and no request has succeeded since it first started, with free retries capped at maxRetries. After a run of new 429s with no success the gate stops retrying and delaying until a request succeeds, so a persistent refusal fails about as fast as before. Without throttling nothing is paced or delayed. Retry counts, the delay for a single request, and the error surfaced when retries run out are unchanged. Known limits: 429s unrelated to the request rate leave the client spacing requests for a while, limits stricter than about 0.5 requests per second are not fully tamed, and a positive Retry-After now pauses the whole client. Verified against simulated servers only, not a live Data Center. Refs #261
Problem
The retry interceptor honors
Retry-Afterliterally, soRetry-After: 0produces a 0 ms delay.Confluence Data Center's rate limiter (a token bucket) sends
Retry-After: 0on 429 responses, and on 200 responses, together withx-ratelimit-limit: 3,x-ratelimit-fillrate: 3andx-ratelimit-interval-seconds: 1. Every retry therefore fired immediately, hit the empty bucket again, and all attempts were used up within milliseconds. On Data Center 9.2.9 (3 requests per second),confluence children <pageId> --recursivesent 36 requests in 186 ms and failed withRate limit exceeded.A negative value, a whitespace-only value and an HTTP-date that is not in the future also resolved to 0 ms.
Fix
retryDelayMshonorsRetry-Afteronly when it names a positive wait: a positive number of seconds or a future HTTP-date, still capped at 60 s. A missing, unparsable, zero, negative or already-past value is treated like a missing header and falls back to the existing exponential backoff (1 s, 2 s, 4 s with the default base delay, capped at 60 s). Positive values behave exactly as before.x-ratelimit-*headers are deliberately not used as a hint. WhileRetry-Afteris positive they add nothing, and the 1 s base backoff already spans the 1 s refill interval reported next toRetry-After: 0.retryBaseDelayMs.Verification
Retry-Afterof"0", numeric0,"00", whitespace, negative seconds, unparsable text and an empty string; past and current HTTP-dates; positive seconds and future HTTP-dates (still honored and capped); the backoff cap; and the interceptor wiring, asserting thesleepdurations for a 429 withRetry-After: 0(1000, 2000) and withRetry-After: 3(3000) without waiting in real time. 8 of them fail onmain(the interceptor test sees[0, 0]); the other 4 guard existing behavior.npm test: 1465 passed (1453 before).npm run lint: clean.Retry-After: 0on every response and rangetAllDescendantPagesover a tree needing 36 requests:mainfails after about 15 ms with 3 of 36 requests served; this branch spaces retries 1 s / 2 s / 4 s and serves 24 before failing.children <pageId> --recursive --max-depth 3gives up after 0.3 s on v2.25.8 and after 7.3 s on this branch: the retries now wait, but the traversal still fails for the reason below.Known limitation
This fixes the retry timing only.
children --recursiveruns up to 10 requests concurrently, which does not fit under 3 requests per second with 3 retries (7 s of waiting). In the same simulation, limiting the traversal to 3 concurrent requests completes the 36-request tree in about 11 s, and raisingmaxRetriesto 10 completes it in about 15 s; jitter alone does not help. Both are behavior changes, so they are left for a follow-up.Type of Change
Testing
Checklist
🤖 Generated with Claude Code