SIGN IN SIGN UP

Deflake replica migration tests (#4684)

The failure in `Sub-replica reports zero repl offset and rank, and fails
to win election - shutdown` test checks for exact roles on exact
replicas for success:

```
wait_for_condition 1000 50 {
    [s -4 role] eq {master} &&
    [s -3 role] eq {slave} &&
    [s -7 role] eq {slave}
}
```

In the failure reported in #4672, failover did happen, but the wrong
replica won (3 rather than 4):

```
s -4 role: slave
s -3 role: master
s -7 role: slave
```

From the logs, replica 4 did start its failover first (`00:26:00.640 *
Myself become the best ranked replica, initiate the election
immediately.`), but its vote request was delayed significantly (most
likely CPU scheduling or network timing), so replica 3's later request
reached the voters first and won epoch 10:

```
24558:M 13 Sep 2026 00:26:01.335 * Failover auth granted to 0e5fe0108fea1b3848216acead193291bb6f0923 (R3) for epoch 10
24558:M 13 Sep 2026 00:26:01.344 # Failover auth denied to 77c1f04b20d64dec79cd058e37ee674f22827cca (R4): already voted for epoch 10
24558:M 13 Sep 2026 00:26:01.344 # Failover auth denied to 1cec4d7bfe9699c731a4d8c480d4d1694bd7061c (R7): already voted for epoch 10
```

Nothing out of the ordinary was noticed when analyzing the code & the
logs, this looks like an unfortunate tail case.

This change sets `debug cluster-failover-delay` to an artificially high
value on the replicas. The freshest replica ignores the delay via the
best ranked replica fast path and elects immediately, while the other
replicas hold off, making the freshest replica's win more deterministic.

Signed-off-by: Bara' Hasheesh <bara.hasheesh@gmail.com>
B
Bara' Hasheesh committed
47c9839b22638e04e93a689fa304ee2985eb5771
Parent: 00a8a19
Committed by GitHub <noreply@github.com> on 9/17/2026, 8:47:29 PM