Fail back to the original cluster¶
Return producers and consumers to the original cluster after a failover, without processing its old backlog twice and without a replication loop.
New in v3.2.0.
Before you start: you failed over from a to b, a is back, and the remote a on b with its user repl-from-b-2m8x4d on a exist, as in Set up disaster recovery.
The target of a remote child must not have remote children of its own, and every link checks this when it is attached and every check_interval_ms after. So the link from b back to a cannot be attached while a's orders still has its link to b, and the order below never has both directions at once. A link that finds its target has a remote child stops in target_has_remote_children without sending.
Fail back¶
-
Prove
ais quiet.rate(narad_messages_produced_total{topic="orders"}[10m])is 0 ona. Records ofordersananode accepted and has not committed yet are invisible to the lag; the detach in step 2 counts them, by node, forordersalone, and refuses while any remain.narad_ingress_dispatch_backlog_recordsis no proof here: it counts every topic the node serves. -
Drain and detach the link from
atob:narad --ctx a topic wait orders orders-dr --lag-zero --stable 60s narad --ctx a topic detach orders orders-dra'sordersnow has no remote children.If you fenced
aby deleting its replicator user onb(failover step 3), the link cannot drain: it stays inauth_failedwith lag above 0,topic waitexits 2 at once (stalled: state auth_failed), and a plain detach answers409(unshipped records). That tail is the stale copy the fence keeps away fromb, so abandon it:narad --ctx a topic detach orders orders-dr --forceThe leader's audit line for
remote_child.deleterecordsforced=trueand what was abandoned (abandoned_lag_messages,abandoned_dispatch_backlog). -
Clear
a's old backlog. Delete and recreatea'sorderswith the same partitions, retention and schema. Everything in it was either processed onabefore the failure or is already onb; without this step,a's consumers would process its backlog a second time. Narad has no purge, so delete and recreate is the tool. Do it only after step 1: anything produced toaafter the detach would be lost. -
Move the topic back with the offload playbook, with the roles swapped (Move a topic to another cluster): stop
b's consumers, attach frombat its consumer frontier, starta's consumers, move the producers toa, provebquiet, and detach.narad --ctx b topic attach orders orders-to-a --remote a --from unconsumed -
Restore disaster recovery. Optionally delete and recreate
b'sordersfirst: otherwise the recordsb's consumers processed age out only afterb's retention, and a failover inside that window processes them again. Then attach the link fromatobas in Set up disaster recovery, step 4. The new attach records the recreated target's new ID.
Both regions producing¶
Two clusters that produce to the same topic and copy it to each other form a cycle, and the loop rule refuses it. What works is one topic per origin: a's orders-a copied to b's orders-a, and b's orders-b copied to a's orders-b, with the consumers in each region reading both. Neither target has remote children, so the rule allows it.
Next steps¶
- Set up disaster recovery: the link, the retention and the alerts.
- Manage remotes: rotate the idle pair of credentials too.