Fail over to the recovery cluster¶
Move producers and consumers to the recovery cluster when the cluster in use, or its region, is lost.
New in v3.2.0.
Before you start: a remote child from the cluster in use, a, to the recovery cluster, b, set up as in Set up disaster recovery, and a's metrics stored outside its region.
Fail over¶
-
Decide that
a's region is lost, not just the link. A link that is down whileastill serves its clients is a lag incident: see Troubleshooting, not this page. -
Record the data at risk: the last stored
narad_fanout_remote_lag_secondsfor the link plus one scrape interval, and the lastnarad_ingress_dispatch_backlog_recordsof eachanode. That is at most whatbmay never receive: the gauge counts every topic the node serves, so for one link it is an upper bound. -
Choose the fence.
- If
a's disks may come back intact, keepa's replicator user onb. Whenareturns, its cursors resume from where they stopped and ship the unshipped tail tob. - If
amay come back as a stale copy, from an old backup or a clone, delete that user onbfirst (narad --ctx b user rm repl-from-a-7f3k9q), so nothingasends can reachb.a's link then never drains: when you fail back, detach it with--forceto abandon the stale tail.
- If
-
Mind a split brain. If
a's region is cut off rather than gone, its producers and consumers may still be running. Stopa's consumers by any path you have. If you cannot, the unshipped tail will be processed in both regions once the regions can talk again. -
Start
b's consumers. They begin atb's oldest retained record, because nothing was ever acked onb: expect duplicates up tob's retention window. -
Move the producers to
b. -
Keep
a's consumers stopped whenareturns, until you fail back.astill holds its backlog from before the failure, whichbhas processed.
When a returns¶
With the fence kept, a's cursors resume from their files and ship the tail: watch narad --ctx a topic children orders until lag_messages reaches 0. Producers that still write to a reach b through the same link. Then fail back, or keep running on b and detach the link once a is quiet.
Next steps¶
- Fail back to the original cluster: return to
awithout processing its backlog twice. - Delivery contract: what each kind of failure can lose.