Skip to content

A small, sturdy message broker. HTTP in, at-least-once out.

When Narad answers 202, your message is already fsynced to disk. It is one Go binary: POST to any node, pull the message, work on it under a lease, ack it. Anything left unacked comes back.

Run it locally

$ NARAD=http://127.0.0.1:7942
$ curl -i -X POST "$NARAD/v1/topics/orders/produce?key=customer-42" \
    -H 'Content-Type: application/json' \
    -d '{"order_id":"ord_123","amount":4999}'
HTTP/1.1 202 Accepted
Date: Mon, 28 Sep 2026 10:49:16 GMT
Content-Length: 0

$ curl "$NARAD/v1/topics/orders/consume?wait=10s"
{"topic":"orders","partition":1,"offset":0,"key":"customer-42",
 "payload":{"order_id":"ord_123","amount":4999},
 "timestamp":1790592556,"receipt_handle":"1:0:8874198393259169482"}

$ curl -i -X POST -H 'Content-Type: application/json' \
    "$NARAD/v1/topics/orders/ack?receipt_handle=1:0:8874198393259169482"
HTTP/1.1 204 No Content
Date: Mon, 28 Sep 2026 10:49:21 GMT
A recorded session against narad server start --dev, after creating the topic orders. The consume response is one line; it is wrapped here to fit.Long lines are wrapped here to fit.

Hit any pod. Narad does the rest.

Put every pod behind one load balancer and send it every produce, consume and ack. Whichever pod catches the request is the right one: your client never looks for a leader, never learns a partition map and never speaks a metadata protocol.

That pod appends the message to its own write-ahead log and fsyncs before it answers 202, so you wait for one local fsync. After the 202 it hands the message to the partition's owner and retries until the owner has fsynced it, read it back and verified it. Only then do consumers see it.

The produce path, step by step

A produce, answered by whichever pod catches it Your service sends POST /produce to the load balancer, which passes it to narad-1; narad-0 or any other pod would have done as well. narad-1 appends the message to its write-ahead log, fsyncs, and answers 202 Accepted, so the producer waited for one fsync. After the 202, in the background, narad-1 hands the message, ord_123, to narad-2, the partition's owner, and retries until narad-2 has fsynced it, read it back and verified it. Narad cluster Your service 1 POST /produce 3 202 Accepted you waited for one fsync Load balancer narad-0 or any other pod narad-1 WAL 2 fsync ord_123 after step 3, in the background narad-2 4 The owner fsyncs, reads back, verifies A produce, answered by whichever pod catches it Your service sends POST /produce to the load balancer, which passes it to narad-1; narad-0 or any other pod would have done as well. narad-1 appends the message to its write-ahead log, fsyncs, and answers 202 Accepted, so the producer waited for one fsync. After the 202, in the background, narad-1 hands the message, ord_123, to narad-2, the partition's owner, and retries until narad-2 has fsynced it, read it back and verified it. Your service 1 POST /produce 3 202 Accepted you waited for one fsync Load balancer narad-0 or any other pod narad-1 WAL 2 fsync ord_123 after step 3, in the background narad-2 4 The owner fsyncs, reads back, verifies Narad cluster
Steps 1 to 3 cost one local fsync. Step 4 happens after the 202 and survives a crash of either pod: the write-ahead log keeps its copy until the owner's copy is verified.

Consume without the ceremony

No consumer groups, no partition assignment, and nothing rebalances when a worker joins. Run one worker or a hundred against the same topic: each message goes to one of them at a time, under a lease that lasts 30 seconds by default. Ack it and it is settled.

If a worker dies mid-job, its lease runs out and the message goes to the next worker that asks. A slow worker extends its lease; one that gives up hands the message back at once. A late ack gets 410 Gone, so you know the work may run twice.

Consuming: leases, acks, extends and nacks

A crashed worker's message comes back Three workers pull from the topic orders, with no consumer group. Worker 1 consumes and acks. Worker 2 takes ord_123 on a 30 second lease and crashes without acking. The lease runs out and ord_123 goes back into the topic. Worker 3, just started, takes it and acks with 204. Topic orders consume, ack worker 1 1 consume, 30 s lease ord_123 3 lease runs out: it comes back worker 2 2 crashed 4 consume and ack worker 3 just started A crashed worker's message comes back Three workers pull from the topic orders, with no consumer group. Worker 1 consumes and acks. Worker 2 takes ord_123 on a 30 second lease and crashes without acking. The lease runs out and ord_123 goes back into the topic. Worker 3, just started, takes it and acks with 204. Topic orders ord_123 1 3 4 worker 1 worker 2 worker 3 2 1 Worker 2 takes it for 30 s 2 Worker 2 crashes, never acks 3 The lease runs out: it comes back 4 Worker 3 takes it and acks
Worker 2 may have done part of the job before it died, so ord_123 can run twice. That is at-least-once: make handlers idempotent.

Built to say yes

Any live node accepts a produce with a local fsync: no leader election and no quorum on the write path, so losing a minority of nodes never stops produces. If a partition's owner is down, the message is committed to a live partition of the same topic instead, and consumers keep consuming.

The price, stated up front: ordering is not guaranteed. Messages already stored on the dead node wait for it to come back, and their partition answers 503 until then. If you need a sequence, carry one in the payload.

The availability trade, in full

A produce while the partition's owner is down Your service sends POST /produce to narad-0, which fsyncs it and answers 202 Accepted. The partition's owner, narad-1, is down, so the hand-off to it is blocked. The message, ord_123, is rerouted to a live partition of the same topic on narad-2, and consumers keep consuming from there. Narad cluster Your service 1 POST /produce 2 202 Accepted narad-0 Owner of the partition narad-1 down ord_123 3 rerouted to narad-2 narad-2 Consumers keep consuming A produce while the partition's owner is down Your service sends POST /produce to narad-0, which fsyncs it and answers 202 Accepted. The partition's owner, narad-1, is down, so the hand-off to it is blocked. The message, ord_123, is rerouted to a live partition of the same topic on narad-2, and consumers keep consuming from there. Your service 1 POST /produce 2 202 Accepted narad-0 narad-1 owner, down ord_123 3 rerouted to narad-2 narad-2 Narad cluster Consumers keep consuming
The 202 from narad-0 depends only on its own disk. It reroutes at once if cluster membership already says the owner is dead, or after 3 seconds of failed hand-offs if it does not. Either way the message lands on a different partition from the messages before it, so order across a failure is not kept.

Deploys like it's nothing

A load balancer, a StatefulSet and a volume per pod: that is the whole architecture. Topics, users and partition owners live in Raft inside the same binary, so there is no ZooKeeper, no BookKeeper and no metadata store to run beside it.

To scale out, raise replicaCount. The new pod joins the cluster and the leader moves partitions onto it. The chart installs from a clone of the repository, once you have created a namespace and one secret:

helm install narad ./charts/narad \
  -n narad --set replicaCount=3 \
  --set image.tag=v3.2.2

The chart's NetworkPolicy fences the Raft port to the Narad pods by default; Raft TLS is the production choice.

Deployment, step by step

Scaling out: a new pod joins and a partition moves onto it Your service sends requests to a load balancer that spreads them over the pods of one StatefulSet, narad-0 to narad-2, each with its own volume. Raft runs inside every pod: narad-0 holds the Raft leader, which replicates to each of the others. A fourth pod, narad-3, is drawn dashed: raising replicaCount adds it, it joins the cluster, and a partition, orders/1, is copied from narad-0 onto it before ownership cuts over. Your service Load balancer narad-0 Raft leader narad-1 Raft narad-2 Raft narad-3 Raft volume volume volume volume Raft: the leader to each follower orders/1 copied, then cut over StatefulSet raise replicaCount Scaling out: a new pod joins and a partition moves onto it Your service sends requests to a load balancer that spreads them over the pods of one StatefulSet, narad-0 to narad-2, each with its own volume. Raft runs inside every pod: narad-0 holds the Raft leader, which replicates to each of the others. A fourth pod, narad-3, is drawn dashed: raising replicaCount adds it, it joins the cluster, and a partition, orders/1, is copied from narad-0 onto it before ownership cuts over. Your service Load balancer StatefulSet raise replicaCount narad-0 Raft leader narad-1 Raft narad-2 Raft narad-3 Raft volume volume volume volume orders/1 copied, then cut over
Raise replicaCount and a new pod, drawn dashed, joins the cluster. The leader then moves partitions onto it, copying each one before it cuts over.

The fine print, up front

What a 202 promises, and what Narad trades for it.

  • A 202 means fsynced to disk. Delivery is at least once, so handlers must be idempotent. Every pull request runs a three-node cluster under load while its nodes restart, and fails unless every message is acked. How the contract is tested
  • Ordering is not guaranteed. Redelivery and rerouting around a dead node both reorder messages. Carry a sequence in the payload if you need one. Every way order breaks
  • Each partition is one copy on one volume. Crashes and restarts lose nothing; a destroyed disk loses that node's partitions. For a second copy, add a replica child or snapshot the volumes. Replication, when you ask for it
  • Fsync costs throughput. On one shared 2 CPU / 2 GB box with 256-byte messages, Narad produced 10,454 msg/s, fifth of six brokers; RabbitMQ's quorum queue, the only other one there that fsyncs before it confirms, was about 1.25 times faster. Batches win most of it back: on three nodes, batches of 100 sustained about 377,000 messages a second through produce, consume and ack. Same compute, measured

Also in the one binary

  • Fan-out children Every message committed to a parent is copied into each child, with its own consumers and retention. Producers change nothing.
  • Replica children A child whose partitions are placed on other nodes than the parent's: an async second copy of a topic, from one API call.
  • Delay children A child with delay_ms receives each message that long after the parent committed it: delayed work with no scheduler.
  • Remote children (v3.2.0) A child whose copy lives on another Narad cluster, in another region if you like: move a topic, or keep a disaster-recovery copy, through the other cluster's own API.
  • Schemas at the broker Give a topic a JSON Schema and a produce that does not fit gets 400 naming the field. It never reaches the log.
  • Any payload Send JSON, text or raw bytes as application/octet-stream. JSON comes back verbatim, text as text, and binary as base64 with a flag that says so.
  • A Go SDK and a CLI The Go client renews leases, retries with jitter and trips per-node circuit breakers, on the standard library alone. The narad binary is the broker and the CLI in one.

Sixty seconds on your laptop

The narad binary is both the broker and the CLI. narad server start --dev runs one node on 127.0.0.1:7942 with auth off, and the Docker command runs the same. That is what the session at the top of this page talked to.

The sixty-second demo · Getting started · Go SDK

docker run --rm -p 127.0.0.1:7942:7942 \
  -v narad-data:/var/lib/narad \
  -e NARAD_SECURITY_ENABLED=false \
  -e NARAD_CLUSTER_ADDR=127.0.0.1:7943 \
  ghcr.io/debanganthakuria/narad:v3.2.2
# builds from source: a minute or more
brew install debanganthakuria/narad/narad
narad server start --dev

Then, in a second terminal, create the topic, and produce, consume and ack it with steps 3 to 5 of the Quickstart, acking with the receipt_handle your own consume returns:

curl http://127.0.0.1:7942/v1/topics \
  -H 'Content-Type: application/json' \
  -d '{"name":"orders"}'