Skip to content

[-] retry failed etcd re-sync until etcd answers, closes #431 - #432

Closed
pashagolub wants to merge 8 commits into
masterfrom
fix-etcd-resync-retry
Closed

pashagolub wants to merge 8 commits into
masterfrom
fix-etcd-resync-retry

Conversation

@pashagolub

Copy link
Copy Markdown
Collaborator

Closes #431

Problem

When the etcd watch is lost, the checker re-arms it and re-reads the leader key once via get(). If that single read fails (e.g. context deadline exceeded while etcd is overloaded), false is sent and the VIP is removed. The re-armed watch is healthy and silent as long as the leader key does not change, and Patroni renews leadership via a lease without rewriting the key, so the VIP never comes back until a restart or a failover. The initial read in GetChangeNotificationStream had the same one-shot flaw.

Change

  • get() now reports whether etcd answered (an absent key counts as an answer).
  • New sync() retries get() every second until etcd answers or the context is done.
  • sync() is used for the startup read and for the re-sync after a lost watch.

The demotion behaviour is unchanged: every failed read still emits false, as before. Only recovery changed.

Tests

  • TestEtcdLeaderChecker_watch_RetriesFailedResync: first watch dead, later watches healthy, first two reads fail; true must come back.
  • TestEtcdLeaderChecker_GetChangeNotificationStream_RetriesFailedInitialGet: initial read fails twice; true must come back.

Both fail on the previous code and pass with the fix; full suite passes.

Note for reviewers

Out of scope, not changed here: a hung etcd (tested by docker pause for 25s) never makes the watch report an error, so the VIP is not removed in that case — neither before nor after this PR. Closing that gap for #354 would need a periodic read next to the watch; worth a separate discussion because it trades against VIP flapping on an overloaded etcd.

A single failed read after a lost watch (or at startup) left the state at
false, and the healthy re-armed watch stays silent while the leader key is
unchanged, so the VIP never came back until a restart.
go.mod requires go 1.26.7, but `go-version: '1.26'` picked the runner's
cached go 1.26.4, and GOTOOLCHAIN=local refused to build.
@coveralls

coveralls commented Sep 16, 2026 •

Copy link
Copy Markdown

Coverage Status

coverage: 91.531% (+0.1%) from 91.434% — fix-etcd-resync-retry into master

@pashagolub pashagolub self-assigned this Sep 25, 2026
@pashagolub pashagolub added the bug label Sep 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Synchronization can race, and the new tests need cancellation-and-wait cleanup.

Review effort: Lite
Findings: None

What changed in this PR

Improves etcd leader-state recovery by retrying failed reads during startup and watch re-synchronization.

Changes:

  • Adds retrying synchronization logic.
  • Adds regression tests for failed reads.
  • Derives the CI Go version from go.mod.
File Description
checker/​etcd_leader_checker.go Implements retrying etcd synchronization.
checker/​etcd_leader_checker_test.go Adds startup and re-sync retry tests.
.github/​workflows/​build.yml Uses the module’s declared Go version.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/build.yml Outdated
GetChangeNotificationStream ran the initial read in its own goroutine next
to the watch. A read that returned the old leader could send false after the
watch had already sent true for a newer leader change, and the VIP stayed
down until the key changed again (#431).

Now the read runs first and the watch starts at the next revision, so every
send comes from one goroutine and later events always win. Watch() blocks
while etcd is unreachable, so reading first also keeps false flowing during
an outage at startup.
The retry tests cancelled their goroutines without waiting, so cleanup could
close the client and the container while those goroutines still ran.
EmitsOnConnectionError looped forever once its context expired because its
break only left the select, so a regression hung the run for 10 minutes
instead of failing.
@pashagolub

Copy link
Copy Markdown
Collaborator Author

Replaced by #436, which polls the leader key instead of watching it and fixes #431 that way.

@pashagolub pashagolub closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Stale State when ETCD had an Issue

3 participants