Move a topic to another cluster¶
Move a busy topic, its producers and its consumers to another Narad cluster without losing a message, using a remote child as the bridge.
New in v3.2.0.
Before you start: every item of Before you start on the source cluster a and the target cluster b, the CLI with a context for each, and a way to stop and start the topic's consumers and to repoint its producers.
The plan: the source keeps taking produces while a remote child copies everything its consumers have not acked yet to the target, the target's consumers take over, and producers move at their own pace. Stragglers still writing to the source reach the target through the link. Once the source is quiet, the link goes.
Steps¶
-
Prepare the target. Create the topic with the source's schema byte for byte, its consumers' and producers' users, and the replicator user (Prepare the target):
narad --ctx a topic schema orders --current > orders.schema.json narad --ctx b topic add orders --partitions 24 --retention 72h \ --schema @orders.schema.json narad --ctx b user add repl-from-a-7f3k9q --grant produce:orders \ --user-password-stdin < repl-passwordLeave out
--schemawhen the source has none. Undo: delete them. -
Register the target as a remote and check it (Register a remote):
narad --ctx a remote add b --url https://narad-b.example.com \ --username repl-from-a-7f3k9q --remote-password-stdin < repl-password narad --ctx a remote ls narad --ctx a remote test b --topic orders --source ordersremote lsmust show every nodeready, with the remote's fingerprint andcredential_version, andremote testmust exit0. Setremotes.allowed_hostsfirst: without it,remote testruns on the one node that took the request, and proves nothing about the other nodes' egress, which step 5's attach checks from every node after the consumers are already stopped. Then deleterepl-password. Undo:narad --ctx a remote rm b. -
Buffer. Raise the source's retention to cover the migration plus the longest target outage you accept:
narad --ctx a topic edit orders --retention 72h. While the link exists, the source's retention cannot go below 24 hours. -
Stop the source's consumers gracefully: they finish what they hold and take nothing new. Producers keep writing to the source.
-
Attach at the consumer frontier. Check first, then attach:
narad --ctx a topic attach orders orders-to-b --remote b \ --from unconsumed --dry-run narad --ctx a topic attach orders orders-to-b --remote b \ --from unconsumedRead the start offsets and any warnings in the dry run, and leave at least 5 seconds between the dry run and the attach: each node checks one remote at most once every 5 seconds, and an attach sent sooner answers
429(retry after itsRetry-After). From now on, everything not yet acked on the source, and everything produced to it later, flows to the target. -
Start the target's consumers. Watch the link with
narad --ctx a topic children orders(running, lag near 0) and the target's consumer lag. -
Move the producers to the target, at any pace. Records the stragglers still send to the source reach the target through the link.
-
Prove the source is quiet. All of these must hold:
rate(narad_messages_produced_total{topic="orders"}[10m])is 0 on every source node.narad --ctx a topic wait orders orders-to-b --lag-zero --stable 60sexits0.narad_fanout_child_dropped_messagesandnarad_fanout_remote_skipped_records_totalfororders-to-bhave not moved since step 5.
Records of
ordersa node answered202for and has not committed yet are invisible to the link's lag. The detach in step 9 counts them, by node, forordersalone (dispatch_backlogin its409). Do not wait fornarad_ingress_dispatch_backlog_recordsto read 0 instead: it counts every topic the node serves, so on a cluster that keeps serving other producers it never does. -
Detach the link, without
--force:narad --ctx a topic detach orders orders-to-b. The detach checks the lag and every node's dispatch backlog oforders; a refusal means something is still unshipped, so go back to step 8. A refusal with onlydispatch_backlogabove 0 usually clears within seconds once nothing produces toorders: detach again. A refusal that names a node innot_answeringis different: that node is down and may hold records only it accepted, so step 8 cannot clear it. Bring the node back, then detach again;--forceabandons whatever its ingress WAL still holds (Troubleshooting). -
Clean up. Delete the replicator user on the target, which also fences any source node that still holds the credential. Then
narad --ctx a remote rm b, and delete the source topic after a grace period.
Why nothing is lost¶
Every source record below the consumer frontier was acked on the source, so it was processed. Every record at or above it is copied with commit before advance, and the target's 202 is the same delivery promise as any produce's. Records produced straight to the target are on the target. Nothing ages out as long as the retention headroom never reaches 0, and the headroom alert pages first.
Duplicates: messages the source's consumers had processed but not acked when they stopped, any message the target stores twice after a resent request, and anything the target's consumers process twice in their own failures. Consumers must be idempotent, as on any Narad topic.
Roll back before step 9¶
- Move the producers back to the source. The target's consumers keep running, and the link keeps feeding them.
- Once the target takes no produces of its own and its consumer lag is 0, stop the target's consumers and start the source's. The source's frontier has not moved since step 4, so they process again what the target already processed: duplicates, no loss.
- Detach the link with
--force:narad --ctx a topic detach orders orders-to-b --force. With the producers writing to the source again, the link always has records in flight, so the plain detach answers409(and429to a retry within 10 seconds). The forced detach abandons only the copies still on their way to the target, which nothing consumes any more; the source keeps every record, and its consumers process them. To detach without--force, stop the producers briefly and detach once the lag is 0, as in step 9 of the steps.
Without stopping the consumers¶
Attach with --from attach while the source's consumers keep running, then wait until they have acked past the link's start on every partition:
narad --ctx a topic wait orders orders-to-b --source-drained
Then start the target's consumers and stop the source's. Duplicates: every record after the attach that the source's consumers processed before the switch. Continue from step 7.
Next steps¶
- Set up disaster recovery: keep a copy on another cluster all the time.
- Replicate a topic to another cluster: every field of the attach and the listing.
- Troubleshooting: a link that does not reach
running.