Configuration reference¶
Look up every setting narad serve reads: its environment variable, its config file key, its default, and the values it accepts.
{
"http": {
"addr": "127.0.0.1:7952"
},
"cluster": {
"addr": "127.0.0.1:7953"
},
"storage": {
"codec": "zstd",
"data_dir": "./narad-data"
}
}
narad serve --config narad.json
The node logs one JSON line per event and serves once a line contains "msg":"http listening". It serves the API on 127.0.0.1:7952, runs Raft on 127.0.0.1:7953, and stores messages compressed with zstd under ./narad-data. Every setting the file leaves out keeps its default.
Under Kubernetes the Helm chart sets most of these for you; which chart value sets which variable is in Helm values reference.
In the tables below, each Setting cell holds the config file key, then the environment variable.
Precedence¶
A setting can come from four layers. Each overrides the one before it:
- the built-in default;
- the JSON config file named by
--config; - environment variables;
- command-line flags of
narad serve.
After the four layers are applied, the whole configuration is checked. Any problem stops the node before it opens any data, with a message that names the setting, such as http.max_consume_wait (20s) must be <= http.shutdown_grace (10s). A key the file may not set stops it the same way: a file with "fsync": "per_write" under storage stops the node with narad: config: config: load file: storage.fsync is an internal setting and cannot be configured.
narad serve takes these flags:
| Flag | Sets |
|---|---|
-- |
the config file to read |
-- |
http. |
-- |
http., as :<n> |
-- |
cluster., as :<n> |
-- |
cluster. |
-- |
storage. |
-- |
log. |
-- |
log. |
-- |
http. |
narad server start --dev is a different entry point for a laptop; it is described in CLI command reference.
HTTP¶
| Setting | Default | Notes |
|---|---|---|
http.NARAD_ |
:7942 |
The client API. Nodes also talk to each other on the same port number over UDP (QUIC). Must differ from cluster.addr. |
http.NARAD_ |
10s |
|
http.NARAD_ |
30s |
Must be longer than http.max_consume_wait, or long polls are cut off mid-wait. |
http.NARAD_ |
60s |
|
http.NARAD_ |
10s |
How long a stopping node lets requests in flight finish. Must be at least http.max_consume_wait. |
http.NARAD_ |
10s |
The longest wait a consume may ask for; longer ones are cut to it. Raising it past 10s means raising http.shutdown_grace too. Set a positive value: 0 passes the checks but falls back to a 30 s ceiling. |
http.NARAD_ |
65536 |
Largest request header block; larger gets 431. At least 4096. |
http.NARAD_ |
4096 |
Open client connections per node; more wait in the listen backlog. 0 removes the cap. |
http.NARAD_ |
1024 |
Concurrent consumes per user, or per client IP with security off, per node; more get 429. A batch consume counts as its max, clamped to the cap. 0 removes the cap. |
http. (v3.1.0)NARAD_ |
0 (off) |
Concurrent produces per user, or per client IP with security off, per node; more get 429. A batch produce counts as its message count, clamped to the cap, and as one while its body is read. v3.0.1 refuses to start with the file key. |
http. (v3.2.0)NARAD_ |
268435456 (256 MiB) |
Batch produce bodies over 1 MiB being read and decoded at once on this node. A body takes its share a megabyte at a time as it grows; one that cannot is answered 503 with Retry-After: 1 and counted in narad_http_batch_body_budget_rejections_total. Bodies of 1 MiB or less never touch it. 0 removes the cap. v3.1.0 refuses to start with the file key. |
http.NARAD_ |
empty | When set, /metrics, /healthz and /readyz are served on this address without credentials, and /metrics leaves the API port. Keep it inside the cluster. May equal http.pprof_addr. |
http.NARAD_ |
false |
Serve /metrics on the API port without credentials. Its series name every topic. Ignored when http.metrics_addr is set. |
http.NARAD_ |
empty | Serves Go's net/http/pprof on this address, without credentials. Keep it inside the cluster. |
Cluster¶
| Setting | Default | Notes |
|---|---|---|
cluster.NARAD_ |
the host name | The node's identity in the cluster. Keep it stable across restarts. |
cluster.NARAD_ |
:7943 |
The Raft transport (TCP). With security on and no Raft TLS files, a node with no peers must bind it to a loopback address such as 127.0.0.1:7943, or set security.allow_plaintext_raft (from v3.1.0). Bind it to loopback only on a node that will never take peers: the address Raft first starts on is recorded in the Raft configuration, the nodes that join later dial that recorded address, and a later cluster.addr does not change it, so a node first started on loopback cannot be grown by rebinding it (Networking and security). |
cluster.NARAD_ |
none | The voters that bootstrap the cluster, the same list on every node. In the environment, id@host:7943,id@host:7943,...; in the file, a list of {"id": ..., "addr": ...}. When set, it lists at least 3 voters. A joining node walks it to find the leader. |
cluster.NARAD_ |
empty | The host:port other nodes dial for this node's Raft transport. Required when the node is not in the peer list; otherwise the node takes the host from its own peer entry. |
cluster.NARAD_ |
empty | Comma-separated IDs of the nodes that may bootstrap a new cluster; every other node joins the existing one. Empty lets every node bootstrap. Never change it after the cluster exists. |
cluster.NARAD_ |
8192 |
Metadata log entries applied since the last snapshot before the next one. Greater than 0. |
cluster.NARAD_ |
120s |
How often the threshold is checked, with up to 2x jitter. At least 5ms. |
cluster.NARAD_ |
10240 |
Log entries kept behind a snapshot. A restarting node further behind than this gets the whole snapshot. |
The three Raft settings are the defaults of the Raft library Narad uses. Leave them alone in production.
Storage¶
| Setting | Default | Notes |
|---|---|---|
storage.NARAD_ |
data |
Everything the node stores: topics/, ingress/ and metastore/. |
storage.config file only |
none |
none or zstd. Compression is off by default. zstd shrinks JSON-like payloads a lot, and more under load, when frames hold more records. |
storage.config file only |
fastest |
zstd level: fastest, default, better or best. Decompression speed does not depend on it. |
storage.config file only |
1800000 (30 min) |
Close a partition log nothing has touched for this long. 0 turns it off; otherwise at least 60000. See Idle partitions. |
storage.config file only |
300000 (5 min) |
How often closed partitions are checked for expired data. 0 turns it off; otherwise at least 60000. |
storage. (v3.1.0)config file only |
1000 |
How long an acked position may wait for a sync to disk. 10 to 60000. v3.0.1 refuses to start with this key. See Consumer offset commit interval. |
storage. (v3.1.0)config file only |
false |
Prepare ingress WAL segments ahead of use. v3.0.1 refuses to start with this key, true or false. See Ingress WAL segment preparation. |
The storage keys in this table are the only ones the config file accepts. The engine's fsync mode, flush and sync cadence and segment size are internal settings with fixed production values, and a config file that sets one is refused (storage.<key> is an internal setting and cannot be configured). What a 202 promises about the disk does not depend on any setting; it is in the delivery contract.
Idle partitions¶
An open partition log holds goroutines, file descriptors and buffers. A node closes any log untouched for storage.idle_log_eviction_ms and reopens it on the next produce, consume or replay.
- Creating a topic opens nothing; it is only a metadata entry until a partition is used.
- Metrics reads never keep a log open, and neither does an attached fan-out child that receives nothing.
- With the cold retention walk off (
storage.cold_retention_walk_msset to0), a log is closed only after retention has finished deleting its expired segments. With the walk on, an idle log is closed whatever its segments, and the walk deletes them once they expire. - Retention deletes data only in open logs. Every
storage.cold_retention_walk_ms, the node opens each closed partition it owns that holds an expired segment, deletes it, and closes the partition again;narad_cold_retention_swept_totalcounts these.
Watch narad_open_partition_logs and narad_idle_logs_evicted_total (Metrics reference). An abandoned topic still keeps its metadata and its last segment on disk until it is deleted.
Consumer offset commit interval¶
New in v3.1.0.
storage.consumer_offset_commit_interval_ms (default 1000, 10 to 60000) sets how long a partition's acked position, and the acks it holds above an unacked message, may wait before they are synced to disk. The node writes them at two cadences:
- Every 100 ms, or every interval when that is shorter, each partition acked since the last tick is written to the operating system's page cache, without a sync. A crash of the broker process loses nothing in the page cache, so it redelivers about the last 100 ms of acks, whatever the interval.
- Once per interval, each partition written since is synced to disk:
fdatasyncon Linux, and on macOSfsyncfollowed by oneF_FULLFSYNCper device for all partitions of the tick. A power loss, a kernel crash or the loss of the machine therefore redelivers up to about the interval plus 100 ms and the sync time: about 1.1 s at the default.
A redelivery is a duplicate, never a loss, and a graceful stop redelivers nothing. v3.0.1 synced every changed partition every 100 ms; set the key to 100 to keep that window, at the cost of more syncs.
The node logs a warning, at most once a minute, when it cannot keep to either cadence: consumer offset commits cannot keep to their interval when one tick takes longer than the tick interval, and consumer offsets wait longer than their durability interval for a device flush when a partition's writes have waited more than twice the interval. Fewer partitions per node, or a faster disk, helps with either.
Ingress WAL segment preparation¶
New in v3.1.0.
storage.ingress_wal_prealloc (default false) makes the ingress WAL create and zero-fill its next 64 MiB segment in the background. Group commits then overwrite blocks that already exist, and their fdatasync does not also have to commit the file's metadata through the file system journal. The gain was measured on ext4 only; APFS showed none. Measure on your own volumes before you rely on it.
- Disk: up to two extra 64 MiB segments per node (the prepared active segment and a spare,
next-segment.prepin<data_dir>/ingress/produce), and one segment of background zero-filling per segment. When the disk is full, preparation fails and the next segment is created the plain way. - Crash recovery: a torn write inside a prepared segment is truncated rather than refused. The details are in Produce path.
- Rollback: turn the setting off and restart each node once on this release before moving to v3.0.1 or earlier. The steps are in Upgrade Narad.
Topic defaults¶
These apply when a topic is created without the field, or with 0. Existing topics keep their values.
| Setting | Default | Notes |
|---|---|---|
topic.NARAD_ |
3 |
At least 3, and at most topic.max_partitions. |
topic.NARAD_ |
108 |
The most partitions a topic may have, at create or later. |
topic.NARAD_ |
604800000 (7 days) |
0 keeps messages forever; any other value is at least 3600000 (1 hour). The Helm chart sets 12 hours. |
topic.NARAD_ |
30000 |
Greater than 0, and no longer than the default retention when that is not 0. |
topic.NARAD_ |
1024 |
Greater than 0. |
topic.NARAD_ |
1024 |
Greater than 0. |
What each topic field does is in Create a topic.
Fan-out¶
| Setting | Default | Notes |
|---|---|---|
fanout.NARAD_ |
4096 |
Most records copied to a child in one batch. |
fanout.NARAD_ |
4194304 (4 MiB) |
Most payload bytes in one batch. |
fanout.NARAD_ |
25 |
How long a partly filled batch waits for more records. |
Larger batches mean fewer syncs on the child and more delay for each record. How the copy works is in Fan-out engine.
Remotes¶
New in v3.2.0.
Where this node's remotes may point, and how much memory remote children may hold. Admins manage remotes through the API; these are the bounds an admin cannot widen through it, so widening one is a config change and a restart. Setting any of them to other than its default with security off stops the node at startup (remotes settings require security.enabled).
| Setting | Default | Notes |
|---|---|---|
remotes.NARAD_ |
empty | Exact host names and *.suffix patterns a remote's URL may name (comma-separated in the environment), matched on the canonical host. An entry that does not canonicalize stops the node at startup. Empty admits any host the address guard allows: a node that holds remotes then logs a warning at startup and exports narad_remotes_allowlist_configured 0. Set it in production (operating condition 4). With it set, remote tests run on every node and report connect times. |
remotes.NARAD_ |
[443] |
Ports a remote's URL may name and a dial may reach. At least one, each 1 to 65535. |
remotes.NARAD_ |
empty | CIDRs the address guard lets through although they are loopback, link-local or metadata addresses, such as 127.0.0.0/8 for a test rig. Logged at startup when set. Leave it empty in production. |
remotes.NARAD_ |
268435456 (256 MiB) |
Records remote children may hold in memory on this node while they wait out a failure; past it, a cursor reads its records again after the wait, skipping those the target already accepted, and counts narad_fanout_remote_rereads_total. At least 0; 0 holds nothing. |
remotes.NARAD_ |
false |
Your attestation that the hop from your ingress to this pod is encrypted (a service mesh with mutual TLS, or an ingress that re-encrypts to the pod). A remote create or password change carries the password in its body, so this node answers those 412 without it. Logged at startup when set. |
What each bound guards is in Remote replication.
Logging and security¶
| Setting | Default | Notes |
|---|---|---|
log.NARAD_ |
info |
debug, info, warn or error. |
log.NARAD_ |
json |
json or text. |
security.NARAD_ |
true |
HTTP Basic authentication and grants on the API, and the shared secret between nodes. |
NARAD_environment only |
generated | The root admin's password, used when a secured cluster first starts with no users. Unset, the node that creates the root admin generates a password and writes it to admin-password in its data directory, mode 0600, and never logs it (from v3.1.0; it used to be logged once). Set on a cluster that already has users, it changes nothing, and a node warns at startup when it is not root's password (from v3.1.0). See Manage the root user. |
NARAD_environment only |
none | The shared secret every node proves to the others on the node-to-node port. Required when security is on and cluster.peers is set. A secured node with no peers and none set generates a random one for the life of the process (from v3.1.0), so no other process can use its node-to-node port; set the same secret on every node, the first one included, before adding peers. Adding peers also needs the first node to have started on a cluster.addr the others can reach, with Raft TLS on every node or security.allow_plaintext_raft: a node whose Raft first started on a loopback address can never take peers (Raft TLS). New in v3.2.0: it also derives the key that seals remote passwords, so sealing one needs at least 32 random bytes as openssl rand -base64 32 or openssl rand -hex 32 prints them, never a passphrase, and a node whose metadata holds a remote refuses to start without one. |
NARAD_ (v3.2.0)environment only |
unset | Set only while you rotate the cluster secret: it opens remote passwords sealed under the previous secret until narad remote reencrypt has run. Cluster RPC never accepts it. Unset it afterwards. See Rotate the cluster secret. |
security.NARAD_security.NARAD_security.NARAD_ |
empty | Mutual TLS for Raft: all three or none. Read once at startup; see Raft TLS certificates. |
security.NARAD_ |
false |
With security on, a node refuses to start without the Raft TLS files unless this says the Raft port is fenced some other way, such as by a NetworkPolicy. A node with no cluster.peers whose cluster.addr is a loopback address needs neither (from v3.1.0; this used to apply only with cluster.peers set). |
security.NARAD_ |
false |
Required to run several nodes with security off, which leaves the API, the node-to-node port and Raft open. One node needs nothing. A node with security off and no cluster secret, alone or not, serves its node-to-node port unauthenticated and logs a warning saying so (from v3.1.0). |
security.NARAD_ |
false |
Also accept the older node-to-node authentication, for a rolling upgrade from a release that used it. Turn it off once every node has rolled. See Upgrade Narad. |
The secrets can only be set in the environment, so config files and ConfigMaps never hold them. Why each setting exists is in Networking and security, and what to set before going live is in the Production checklist.
Config file¶
The file named by --config is JSON. It is strict:
- A key the loader does not know, at any level, stops the node from starting. That includes a key from a newer release, which matters when rolling back: remove
storage.consumer_offset_commit_interval_ms,storage.ingress_wal_preallocandhttp.max_produce_in_flight_per_identityfrom the file before a node runs v3.0.1 or earlier (Upgrade Narad). v3.1.0 likewise refuseshttp.max_batch_body_bytes_in_flightand theremotesblock, both new in v3.2.0. - Durations are strings with a unit, such as
"10s"or"500ms". A bare number is refused. - JSON has no comments.
A file that sets a few common values:
{
"http": {
"addr": ":7942",
"max_consume_wait": "10s",
"metrics_addr": ":9100"
},
"storage": {
"data_dir": "/var/lib/narad",
"codec": "zstd",
"compression_level": "fastest"
},
"topic": {
"default_retention_age_ms": 43200000
},
"log": {
"level": "info",
"format": "json"
}
}
The Helm chart writes narad.config from its values into this file (Helm values reference).
Go runtime¶
New in v3.1.0.
On Linux, when GOMEMLIMIT is unset, narad serve sets Go's soft memory limit to 90% of the process's cgroup memory limit, so the garbage collector works harder near the limit instead of letting a burst get the process killed. It logs go memory limit set from the cgroup memory limit (set GOMEMLIMIT to override) once at startup. Any non-empty GOMEMLIMIT, off included, wins; an empty one counts as unset. Without a cgroup memory limit, nothing is set. The binary never changes GOGC.
Tuning¶
| You want | Change |
|---|---|
| Less disk | storage. |
| Longer long polls | http., with http.shutdown_grace at least as long and http.write_timeout longer |
| Larger fan-out batches on slow disks | a higher fanout.linger_ms |
| Fewer offset syncs under heavy ack traffic (v3.1.0) | a higher storage.consumer_offset_commit_interval_ms; a power loss then redelivers more acked messages |
| Fewer redeliveries after a power loss (v3.1.0) | a lower storage.consumer_offset_commit_interval_ms; 100 matches v3.0.1 |
| A ceiling on one user's concurrent produces (v3.1.0) | http. |