When the etcd watch dies, watch() re-arms it with elc.Watch(...) and only then calls elc.sync(ctx, out) (etcd_leader_checker.go#L185-L191).
clientv3.Watch() does not return until the gRPC watch stream is open. It waits on wr.retc (client/v3/watch.go:376 in v3.7.1). While etcd is unreachable, the re-arm call blocks until etcd answers or the context is cancelled, so sync never runs during the outage and the checker sends nothing. The last state stays in place, so a node that held the VIP keeps it for the whole outage.
TestEtcdLeaderChecker_GetChangeNotificationStream_EmitsOnConnectionError expects false while etcd is unreachable. That test only covers startup, where #432 now runs sync before Watch(). The reconnect path has no such test.
This also limits the retry from #431. sync only starts after Watch() has reconnected, so in practice it retries only when the watch works and Get fails.
Observed while working on #432. Moving the first sync after Watch() at startup made EmitsOnConnectionError receive no value at all against an unused port.
Possible fix: in the reconnect path, run sync before re-arming the watch and watch from rev+1, the same order #432 uses at startup. sync then sends false every second until etcd answers.
When the etcd watch dies,
watch()re-arms it withelc.Watch(...)and only then callselc.sync(ctx, out)(etcd_leader_checker.go#L185-L191).clientv3.Watch()does not return until the gRPC watch stream is open. It waits onwr.retc(client/v3/watch.go:376in v3.7.1). While etcd is unreachable, the re-arm call blocks until etcd answers or the context is cancelled, sosyncnever runs during the outage and the checker sends nothing. The last state stays in place, so a node that held the VIP keeps it for the whole outage.TestEtcdLeaderChecker_GetChangeNotificationStream_EmitsOnConnectionErrorexpectsfalsewhile etcd is unreachable. That test only covers startup, where #432 now runssyncbeforeWatch(). The reconnect path has no such test.This also limits the retry from #431.
synconly starts afterWatch()has reconnected, so in practice it retries only when the watch works andGetfails.Observed while working on #432. Moving the first
syncafterWatch()at startup madeEmitsOnConnectionErrorreceive no value at all against an unused port.Possible fix: in the reconnect path, run
syncbefore re-arming the watch and watch fromrev+1, the same order #432 uses at startup.syncthen sendsfalseevery second until etcd answers.