Skip to content

etcd checker keeps the VIP while etcd is unreachable #435

Description

@pashagolub

When the etcd watch dies, watch() re-arms it with elc.Watch(...) and only then calls elc.sync(ctx, out) (etcd_leader_checker.go#L185-L191).

clientv3.Watch() does not return until the gRPC watch stream is open. It waits on wr.retc (client/v3/watch.go:376 in v3.7.1). While etcd is unreachable, the re-arm call blocks until etcd answers or the context is cancelled, so sync never runs during the outage and the checker sends nothing. The last state stays in place, so a node that held the VIP keeps it for the whole outage.

TestEtcdLeaderChecker_GetChangeNotificationStream_EmitsOnConnectionError expects false while etcd is unreachable. That test only covers startup, where #432 now runs sync before Watch(). The reconnect path has no such test.

This also limits the retry from #431. sync only starts after Watch() has reconnected, so in practice it retries only when the watch works and Get fails.

Observed while working on #432. Moving the first sync after Watch() at startup made EmitsOnConnectionError receive no value at all against an unused port.

Possible fix: in the reconnect path, run sync before re-arming the watch and watch from rev+1, the same order #432 uses at startup. sync then sends false every second until etcd answers.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions