Upgrade Narad¶
Move a cluster to a new Narad release with a rolling restart, and roll it back safely if you need to.
Before you start: every pod ready, Helm access to the release, and the changelog entries for every release between yours and the target.
Upgrade¶
-
Get the chart of the target release. In a clone made as in Deploy on Kubernetes:
git fetch --depth 1 origin tag v3.2.2 git checkout v3.2.2 -
Read the version notes for every release you cross.
-
Upgrade, pinning the new image tag:
helm upgrade narad ./charts/narad -n narad \ --reset-then-reuse-values \ --set image.tag=v3.2.2 kubectl rollout status statefulset/narad -n narad--reset-then-reuse-values(Helm 3.14 or later) keeps your settings and gives any value the new chart adds its default. If you keep a values file, changeimage.tagthere and pass it with-finstead.
The StatefulSet restarts one pod at a time, from the highest-numbered pod down, and waits for each to be ready before it restarts the next. A leader that stops hands leadership to another node first and logs leadership transferred before shutdown. While a pod restarts, the other nodes keep accepting produces. Records meant for the restarting pod's partitions wait in the accepting node's ingress WAL, and go to other partitions of the topic if the pod stays away for more than a few seconds. Consumers wait for that pod's partitions, and some messages can be delivered twice. The failure matrix has the details for a rolling restart.
What is stable across releases, and what the docs describe, is in API stability and versions.
Roll back¶
A rollback is the same upgrade with the older tag, using the chart of the release you go back to:
git fetch --depth 1 origin tag vX.Y.Z
git checkout vX.Y.Z
helm upgrade narad ./charts/narad -n narad \
--reset-then-reuse-values \
--set image.tag=vX.Y.Z
kubectl rollout status statefulset/narad -n narad
vX.Y.Zis the release you go back to, in all three places.- Do not pick v3.0.0: a v3.0.0 node that has been up for more than 24 hours cannot be reached by its peers (version notes).
Prefer this to helm rollback. A helm rollback to a revision with a smaller replicaCount is a scale-in that skips the decommission (Scale out and in). helm rollback renders no templates, so only a hook can stop it: the scale-in guard (from v3.1.0), a pre-rollback hook Job of the revision you roll back to, refuses while a pod being deleted is still a cluster member. A revision rendered by an older chart, v3.0.1's included, has no guard and refuses nothing.
Leave a build newer than v3.0.1¶
Releases from v3.1.0 on write some data and accept some settings that v3.0.1 and earlier do not understand. Before a node goes back to v3.0.1 or earlier:
- Drain the ingress WAL. Newer builds write accepted produces in a record format older binaries cannot read; a rolled-back node stops delivering from the first such record it still holds, while it keeps answering
202. Pause producers, keep every partition owner up, and wait untilnarad_ingress_dispatch_backlog_recordsis 0 on every node, in a sample taken after the pause (the gauge refreshes every 5 seconds). Then roll back and resume. - Turn off segment preparation, if you turned it on. Remove
storage.ingress_wal_preallocfromnarad.config, restart every pod once on the newer image (kubectl rollout restart statefulset/narad -n narad), then roll back. Skipping this can leave an older binary unable to start after a crash, with acorrupt frameerror. - Remove the new config keys. Take
storage.consumer_offset_commit_interval_ms,storage.ingress_wal_preallocandhttp.max_produce_in_flight_per_identityout ofnarad.config, even where they hold their defaults. An older binary with one of them in its config file stops at startup withstorage.ingress_wal_prealloc is an internal setting and cannot be configured, orjson: unknown field "max_produce_in_flight_per_identity", before it opens any data. The environment variableNARAD_HTTP_MAX_PRODUCE_IN_FLIGHT_PER_IDENTITYcan stay. - Move clients off the batch forms. An older node answers a batch produce
404, ignoresmaxon a consume and returns one message in the single shape, and answers a batch ack400(receipt_handle required). - Set
GOMEMLIMITif you relied on the derived one. Older binaries do not set it from the pod's memory limit; add it toextraEnv. -
Stop every pod cleanly, and restart a crashed one on the newer image first. Give pods a
terminationGracePeriodSecondslong enough to finish their shutdown: after the HTTP drain (up tohttp.shutdown_grace, 10 s by default), a clean stop syncs the consumer offsets of every partition acked since their last sync, then writes and syncs thehwmfile of every partition written to since it was opened, one partition after another. Measured on local disk under Linux (in a container), that is about 0.1 to 0.4 ms and 0.6 to 1.6 ms per partition, so the chart's default 30 s covers roughly 10,000 such partitions; a network volume syncs slower, so measure on yours. A pod that crashed, or was killed at the end of its grace period, can leave a partition'shwmfile empty, and an older binary handles an empty file badly in two ways:- It reads the file as "take the record tail" when it opens the partition's log, but as a boundary of 0 when it reads a closed partition, and rewrites the file only at its first commit there. On v3.0.1 and earlier the transfer listing is one of those readers. If the rolled-back node is the source of a move that force-promotes (the source dies after a catch-up listing and stays dead past
ForcePromoteAfter), the new owner can get a copy that hides every record until its next commit and then redelivers the whole partition, or a partial copy. The returning source quarantines its own copy, so those records can be recovered by hand. Live moves and decommissions are not affected: their freeze opens the log. - It takes the boundary from the record tail without syncing the segment first. A power loss before its first commit on the partition can take back records it already served, and once their acks persist, their offsets are reused and records committed there later can be skipped.
So after any unclean stop, start the pod once on the newer image and stop it cleanly before you roll it back: startup opens every partition the pod owns, and a clean stop writes each one's exact boundary. 7. Roll back only while a member still runs v3.0.x. A release that adds a Raft entry type uses it only once every member reports a release that knows it (Raft entry types and upgrades). That moment is the rollback boundary: the last member to upgrade starts on this release and its first heartbeat reaches the leader. Before it, a rollback is safe: the member still on v3.0.x holds every new type back while you roll the others back. After it, a rollback to v3.0.x is unsupported, even if the leader has not used a new type yet. A node you stop to roll it back keeps counting as this release for as long as it is down, because its member record holds the last report it sent, so the leader can use a new type in that time: the dead mark it writes for that node after 30 s of silence is one, and so is a topic create. v3.0.x skips such an entry silently, so a topic created in that time can be missing on that node, and on the others once they are rolled back too; later releases stop applying instead (
metastore: stopped applying raft entries). To see where the roll stands, check the image every pod runs (kubectl get pods -n narad -o custom-columns=NAME:.metadata.name,IMAGE:.spec.containers[0].image) or theversionon each pod'snarad serve startingline;narad cluster membersdoes not show a member's release. While a member holds the new types back, the leader names it and its build in thereasonofmetastore: not using a new raft entry type yet, logged whenever a write would have used one. Neither that line's absence nor the absence of the first-use line (metastore: every member applies raft entry type) makes a rollback safe once every pod runs this release. A node that rolls back and also loses its volume is refused when it rejoins a cluster whose every member runs the newer release (409,older_release). v3.1.0 adds entry types 23 to 32 (the list). - It reads the file as "take the record tail" when it opens the partition's log, but as a boundary of 0 when it reads a closed partition, and rewrites the file only at its first commit there. On v3.0.1 and earlier the transfer listing is one of those readers. If the rolled-back node is the source of a move that force-promotes (the source dies after a catch-up listing and stays dead past
Partition segments, consumer position files and fan-out cursor files need nothing: each binary reads what the other wrote. Neither does fsm.db: v3.0.1 ignores the fsm_meta bucket newer builds add, and a newer build that finds the file written by v3.0.1 rebuilds it or restores a snapshot (Restarts). An empty hwm file left by a crash hides no acked record from an older binary once it opens the log; step 6 covers what such a file costs on a closed partition. One gap stays open: a message stored without a key and fanned out to a replica child after the rollback is placed round-robin, so both copies can land on one node. The changelog's "Upgrade and rollback notes" have every detail.
Leave a build newer than v3.1.0¶
New in v3.2.0.
- Never after a remote write. v3.2.0 adds Raft entry types 33 to 37 for remotes. Once a remote has been created or a remote child attached, a rollback to v3.1.0 or earlier is unsupported, even after every remote and remote child is deleted: v3.1.0 stops applying at the first such entry, and refuses at startup a database that applied one (
written by a newer Narad release). Before the first remote write a rollback is fine, but make no remote write while any node is being rolled back: a stopped node keeps counting as v3.2.0 (step 7 above), so the leader would accept the write, and the node would stop applying when it starts on v3.1.0. - Remove the new config keys. v3.1.0 refuses to start with
http.max_batch_body_bytes_in_flightor anyremotes.*key innarad.config. The chart passes itsremotesvalues, and thecluster-secret-previouskey, as environment variables, which v3.1.0 ignores. - Move producers back to v3.1.0's batch limits. v3.1.0 answers a batch of more than 100 messages
400, a body over 1 MiB413, and a compressed body400.
Version notes¶
Read the notes for each release boundary you cross, in either direction.
- To v3.2.2. No stored format or config key changes, so it rolls onto, and back to, v3.2.1 node by node. What clients can see:
- Forwarded acks say whether they landed. An ack, extend or nack that never left the node you called answers
503withRetry-After: 1(not applied, retry it);502now means only that the owner may have applied it. Until every node runs this release, a node on v3.2.1 still answers502for both, so keep treating502as possibly applied. - Retry-After on owner-down. The
503for an ack, extend, nack, replay or pinned consume whose partition owner is down carriesRetry-After: 1. - Remote writes answer once applied. A remote write answered by a follower returns
2xxonly after that node applied it; otherwise503withRetry-After: 2, which says the change is committed and must not be sent again (Status codes). - Blind check reports (no
remotes.allowed_hosts) leave outtarget_serves_ids, and a blind attach no longer adds warnings drawn from the target's answers.
- Forwarded acks say whether they landed. An ack, extend or nack that never left the node you called answers
- To v3.2.1. Nothing to do in either direction: v3.2.1 changes no stored format, config key or API, so it rolls onto, and back to, v3.2.0 node by node.
-
To v3.2.0.
- Remotes and remote children need every member upgraded. No action during the roll. Remote writes and remote child attaches answer
412, naming the member, until every member, dead ones and Raft servers without a member record included, runs this release with security on and legacy cluster authentication off (Manage remotes). The rollback boundary for their Raft entry types is the first remote write (above). - Batch produce takes more. A node on this release takes up to 1,000 messages in a body of up to 16 MiB, each payload at most 1 MiB, optionally zstd or gzip compressed, and answers a body over 1 MiB
503when its budget for large bodies is full (http.max_batch_body_bytes_in_flight). Until every node a client can reach runs this release, keep batches within v3.1.0's limits. - New refusals clients can see.
409for a produce, consume or ack on a remote child's stub, for a change to a stub, for a retention below 24 hours on a parent with remote children, and for a delete or detach while a remote child's records are unshipped;412and429on the remote routes (Status codes). - Deletes during the roll. A delete or detach that involves a remote child travels to the leader as a remote write; while the leader still runs v3.1.0 there are no remote children to involve, and every other delete keeps v3.1.0's path.
- Remotes and remote children need every member upgraded. No action during the roll. Remote writes and remote child attaches answer
-
To v3.1.0.
- Single nodes with security on close their node RPC plane to other processes. No action. A node with no peers and no
NARAD_CLUSTER_SECRETgenerates a secret for the life of the process, so only it can use its node-to-node port (Networking and security). Before you grow such a node into a cluster, set the sameNARAD_CLUSTER_SECRETon it and restart it; a joiner cannot authenticate to it otherwise. It also needs acluster.addrthe other nodes can reach, and either the Raft TLS files on every node orNARAD_SECURITY_ALLOW_PLAINTEXT_RAFT=truewith 7943/tcp fenced: a node with no peers configured never runs the join loop, so one the others cannot reach over Raft stays cut off from their Raft. A node first started on a loopbackcluster.addrcannot be grown this way: the Raft configuration keeps the address its Raft first started on, and a latercluster.addrdoes not change it. Start a new cluster whose first node starts on an address the others can reach, and move the workload to it (Networking and security). Nothing is persisted, so a rollback needs no step. With security off the plane stays open, and startup now says so in a warning. - Secured single nodes need Raft TLS off loopback. A node with security on and no peers, whose
cluster.addris not a loopback address, now refuses to start without the Raft TLS files orsecurity.allow_plaintext_raft, as a node with peers already did. Before upgrading such a node, set the TLS files on every node of the cluster, bindNARAD_CLUSTER_ADDRto127.0.0.1:7943only on a node that will never take peers (a node whose Raft first starts on loopback keeps that address in the Raft configuration, so it cannot be grown by rebinding it later), or setNARAD_SECURITY_ALLOW_PLAINTEXT_RAFT=trueonce 7943/tcp is fenced (Networking and security). The quickstart'sdocker run,--devand Helm installs are unaffected. - Metadata restarts. No action. Each node now records in
fsm.dbwhich Raft entries it holds, and a restart applies only the ones it lacks. The first restart on this release of a node with no Raft snapshot yet (fewer thancluster.raft_snapshot_thresholdmetadata changes, 8192 by default, since the cluster started) rebuildsfsm.dbfrom the Raft log and keeps the old file asfsm.db.stale: delete it once the pod is ready (Troubleshooting). A node rolled back to v3.0.1 and upgraded again does the same, or restores its snapshot. - A node whose metastore cannot write stops. A full data volume or an I/O error on a metadata write now stops the node after 30 s of retries, and it restarts in a loop until the volume is fixed (Troubleshooting); v3.0.1 dropped the change on that node and stayed ready. Keep free space on the data volume and alert on it: pods that share a full volume all stop.
- Heartbeats and joins report the release. No action. Each node's heartbeat now carries its build and the newest Raft entry type it applies, and a join request carries the latter, so the leader can hold a new entry type back until every member knows it (Raft entry types and upgrades). A v3.0.x node refuses those requests with a
400, and the upgraded node resends them at once without the new fields, so no heartbeat is missed. A node that applies fewer Raft entry types than every member is refused when it joins (409, codeolder_release): upgrade it to the cluster's release (Troubleshooting). For rollbacks, see step 7 of Leave a build newer than v3.0.1. - New Raft entry types. No action during the roll, but once the last pod starts on this release a rollback to v3.0.x is unsupported (step 7 above). Once every member, dead members and Raft servers without a member record included, reports this release, the leader starts using the entry types 23 to 32 for topic creates, topic changes, placement, the orphan assignment row prune, dead marks and user deletes, and logs
metastore: every member applies raft entry type <n>; using <name>at info the first time it uses each. Until then it writes the entries v3.0.x applies and, whenever a write would have used a new type, logsmetastore: not using a new raft entry type yet; proposing the entries every member appliesat info, at most once a minute per type, naming the member holding it back and its build inreason. The first-use line is not the rollback boundary: the boundary is the moment every member has reported this release, and a rollback after it is unsupported even if no new type has been used yet, because a node stopped for a rollback still counts as this release while it is down. Clients can see one new answer:409topic changed since it was read(Status codes). - Partition-move answers change. No action.
CompleteMoveanswers409for a refused flip and404for a partition with no assignment (was503), andGetAssignmentis answered only by the Raft leader. While the leader still runs an older release, a move whose flip fails waits with its copy installed until the leader is upgraded or the move is re-planned; no data is at risk (Rebalance). When the move is re-planned, the worker ends and leaves the install at the partition's path, and the destination's stale-copy sweep quarantines it (renamed to<partition>.quarantineand logged at error level) instead of reclaiming it. Its records are usually also on the owner it was copied from; check the owner, then decide (Troubleshooting). - The chart fences Raft with its own NetworkPolicy.
networkPolicy.enableddefaults totrueandsecurity.allowPlaintextRafttofalse. An upgrade with--reset-then-reuse-values, as in Upgrade, takes both defaults: the release gains a NetworkPolicy that admits 7943/tcp and 7942/udp only from its own pods and leaves the API and metrics ports open. It needs a CNI that enforces NetworkPolicy; on one that does not, the ports stay as open as before. To upgrade without the policy, passnetworkPolicy.enabled=falsewithsecurity.clusterTLS.enabled=trueorsecurity.allowPlaintextRaft=true; with neither, the render fails and names the three fixes (Production checklist). An upgrade with--reuse-valueskeeps the oldallowPlaintextRaft: trueand renders no policy, as before: check that something fences the port, or passnetworkPolicy.enabled=true. The nodes themselves need nothing. - Scale-in is approved per size, and checked on rollback.
allowScaleInis no longer read: after decommissioning, setallowScaleInToto the newreplicaCount. A hook Job, the scale-in guard, runs before everyhelm upgradeandhelm rollbackand refuses one that would delete a pod still listed bynarad cluster members(Scale in). On a scale-in it needs the security secret'sadmin-passwordkey to hold root's current password. A rollback to a revision rendered by the v3.0.1 chart runs no guard.
- Single nodes with security on close their node RPC plane to other processes. No action. A node with no peers and no
- From v3.0.0. Roll every node. A v3.0.0 node that has been up for more than 24 hours cannot be reached by its peers for forwarded produce, consume and ack until it restarts. v3.0.1 fixes this.
- Across v2.2.0 (node-to-node authentication). Nodes before and after v2.2.0 cannot talk over the node RPC plane. Roll twice: first with
--set security.allowLegacyClusterAuth=true, so upgraded pods still speak the old protocol to the pods not yet rolled, then, once every pod runs the new image, with it back atfalse. If you skip the first roll, forwarded requests between old and new pods fail until the roll finishes. - Across v2.2.0 (Raft TLS). From v2.2.0, a secured multi-node cluster refuses to start unless Raft runs over TLS or
allowPlaintextRaftsays the port is fenced. Up to v3.0.1 the chart setsallowPlaintextRaft: trueby default; from v3.1.0 it says the port is fenced only when its NetworkPolicy or yourallowPlaintextRaftdoes (see the v3.1.0 notes above). See the Production checklist. - Across v2.2.0 (file modes). From v2.2.0, Narad creates its data directories with mode
0700and its files with0600. It sets the mode only when it creates a file, so files an older release created keep0755and0644. On a host where other users can read the data directory, tighten those by hand. - Config files. Every release since v1.0.0 refuses a config file key it does not know, and fails at startup. When you roll back, remove keys the older release does not have first (Configuration reference).
replicaCountin a rollback. Going back to fewer replicas is a scale-in. Decommission first (Scale out and in).
Next steps¶
- Changelog: every change and upgrade note, per release.
- API stability and versions: what can change in
/v1, and which release these docs describe. - Monitor and alert: watch the cluster through the roll.