Skip to content

[!] poll the etcd leader key instead of watching it - #436

Merged
pashagolub merged 1 commit into
masterfrom
etcd-poll
Sep 28, 2026
Merged

pashagolub merged 1 commit into
masterfrom
etcd-poll

Conversation

@pashagolub

Copy link
Copy Markdown
Collaborator

The etcd checker now reads the leader key once per interval instead of watching it, the same way the Consul and Patroni checkers work. Closes #431, closes #435, and replaces #432.

While etcd is completely unreachable, clientv3 keeps retrying the watch and never closes the watch channel. With a real etcd container stopped under an open WithRequireLeader watch, nothing arrived on the channel for 40 seconds. So the checker sent nothing, and the node kept its VIP for the whole outage (#435). Every earlier watch fix (#228, #394, #431, #432) handled one more way the watch can go quiet or send stale state.

Changes:

  • GetChangeNotificationStream calls get in a loop and waits interval ms between reads. A failed read sends false, and the next successful read sends the current value.
  • "Current value from DCS" is logged only when the value changes, since it is now read every second. Read failures still go through the existing log throttler.
  • interval affects etcd again, so the README note "Doesn't affect etcd checker since v2.3.0" is removed. behaviour_test.sh drops --interval 3000, because with polling its sleep 2 needs the default 1000.
  • The watch-only tests are removed. There are two new tests. FollowsLeaderKey checks that changes to the key are picked up. Outage checks that reads failing sends false and that recovery sends true again; it fails against the old watch code. EmitsOnConnectionError no longer hangs for 10 minutes when it fails.

Trade-off: a leader change is noticed up to interval ms later (1s by default) instead of immediately.

The watch kept producing bugs: cancelled watches (#228), events lost while
the watch was down (#394), a failed re-sync leaving a healthy watch silent
(#431), and the startup read racing the watch. While etcd is unreachable the
client retries the watch without closing its channel, so the checker sent
nothing and a node kept its VIP for the whole outage (#435).

The etcd checker now reads the key once per interval, like the Consul and
Patroni checkers. A failed read sends false, and the next successful read
sends the current value. The value is logged only when it changes, and read
failures stay throttled. interval applies to etcd again, so the README note
saying otherwise is gone and the e2e test uses the default interval.
@coveralls

Copy link
Copy Markdown

Coverage Status

coverage: 91.383% (-0.05%) from 91.434% — etcd-poll into master

@pashagolub
pashagolub merged commit 752aef8 into master Sep 28, 2026
4 of 5 checks passed
@pashagolub
pashagolub deleted the etcd-poll branch September 28, 2026 11:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

etcd checker keeps the VIP while etcd is unreachable Stale State when ETCD had an Issue

3 participants