Cluster
Nagare runs fine on a single instance, but most production deployments will scale out. This page explains the two pieces that make multi-instance deployments work safely:
- Distributed locks stop two instances from running the same outbox runner or singleton subscription at the same time.
- The lease registry is a small visibility table the dashboard reads so you can see which node holds which lock.
The locks are part of the storage backend you already configured. The lease registry is opt-in — turn it on and the Cluster page in the dashboard fills in.
Why locks at all
A few things in Nagare are meant to run as exactly one instance:
- The outbox runner. If two runners drain the same
nagare_outboxtable they will fight over rows, double-dispatch, and burn retry budget. - Singleton-flavoured subscriptions — the catch-up path of a
Subscriptionthat owns a checkpoint, andEventRouteSubscriptionfor process managers. Two of either running side-by-side would step on the same checkpoint row and deliver each event twice.
Live taps don't need a lock — they hold no state. Per-aggregate work is already serialised by Orleans grain placement.
What you get out of the box
When you call AddPostgresOutbox, AddMySqlOutbox, or AddSqlServerOutbox, the package also registers an ILockProvider keyed to that database:
| Backend | Implementation |
|---|---|
| Postgres | PostgresDistributedLock from Medallion.Threading (pg_try_advisory_lock) |
| MySQL 8 | GET_LOCK |
| SQL Server | sp_getapplock with session scope |
| SQLite | No-op — single-process by definition |
The runners and subscriptions ask the provider for a named lock when they start. If they get a handle they own that responsibility until the handle is disposed or the connection drops; if they don't, they sleep and try again on the next tick. The outbox and timeout runners keep their lock across ticks and release it on shutdown, when the lock is lost, or when a tick fails.
That part already works. You don't need to wire anything for safety. But unless you turn the lease registry on, the dashboard can't tell you who is holding what — the locks live entirely inside the database engine.
Turning on the lease registry
Add one line after your storage registration (AddNagarePostgresStorage, AddNagareMySqlStorage, AddNagareSqlServerStorage, AddNagareSqliteStorage):
builder.Services.AddNagare();
builder.Services.AddNagarePostgresStorage(builder.Configuration);
builder.Services.AddRelationalLeaseRegistry();The order matters. The storage registration adds the real ILockProvider, and the registry decorates the lock provider registered when it is called; a lock provider registered later replaces the decorated one, so locks would be taken without writing leases. A stream pipeline that finds a configured lease registry behind a lock provider that does not record leases logs a warning naming this fix at start.
That does three things:
- Registers
RelationalLeaseRegistry, backed by a singlenagare_leasestable in the sameDbConnectionyour stores already use. Same row format on Postgres, MySQL, SQL Server, and SQLite. - Wraps your
ILockProviderwithRecordingLockProvider. The wrapper is transparent — it forwards everything to the inner provider — but on each successful acquire it writes a row tonagare_leases, stampsrenewed_atevery 30 s while the lock is held, and once more on dispose. - Adds a hosted service that creates the table on startup and sweeps rows whose
renewed_atis older than 2 minutes — the leftovers of a node that crashed hard and never got to release. Long-held leases survive the sweep because their holders keep renewing.
There is no separate Nagare.Cluster package. The interfaces live in Nagare.Cluster, the relational implementation in core Nagare. The decoration is part of the same DI call.
Schema
CREATE TABLE IF NOT EXISTS nagare_leases (
lock_name TEXT PRIMARY KEY,
node_id TEXT NOT NULL, -- "<machine>/<pid>"
machine TEXT NOT NULL,
pid INTEGER NOT NULL,
kind TEXT, -- "outbox" | "timeouts" | "subscription" | null
acquired_at TEXT NOT NULL, -- ISO-8601, UTC
renewed_at TEXT NOT NULL
);Lock names follow conventions: outbox locks are prefixed nagare-outbox-, timeout locks nagare-timeouts-; subscription and event-route locks are the catchpoint key (book-catalog, event-route-InterLibraryLoan-BookEvent). The default classifier maps the prefix to the kind column. If you have your own scheme, pass a classifier to AddRelationalLeaseRegistry.
What the dashboard shows
/_nagare/cluster lists three things:
- The local node — machine name, PID, uptime. This is whatever you're connected to.
- Active leases — every row in
nagare_leasesfrom any node, with the holder's machine/pid and how recently the row was renewed. The current node's holdings are taggedthis node. - Background runners —
IHostedServices registered on the local node and their state.
In a multi-instance deployment, every node sees the same lease list (because they all read the same table) but only its own runner list. That's the right shape for "is the outbox actually running somewhere, and who?"
Behaviour you should know about
- Release stamps
renewed_at; it does not delete the row. A released lease stays visible until the next node boot sweeps rows older than 2 minutes, so the page shows who held a lock recently as well as who holds it now. Long-held leases (pipelines, runner leadership) renew every 30 s and survive the sweep; a renewal also recreates a row that a sweep removed while its holder was still alive. - Leaders step down on failure. A runner whose tick throws releases its lock before backing off, so leadership moves within one
IdleDelay. Per-row transient dispatch failures do not count — they follow the retry ladder on the current leader. - Lost-lock detection is not instant. After an unclean lock loss the old holder learns of it within about 60 s; see the outbox guide's limits for what that means.
- Registry failures never block correctness. Every write the wrapper does to
nagare_leasesis in a try/catch with a warning log. If the table is missing or the connection is broken, the inner lock still works — you just lose visibility. - Acquire failures don't leave rows. If two nodes fight for the same lock, only the winner ever calls
RecordAcquired. The loser getsnullfrom the inner provider and bails before reaching the registry. - Standby health reads it. A stream pipeline on standby is only reported healthy while the lock holder keeps renewing its row; a holder whose row goes stale for
StandbyHolderStaleAfter(default 90s) makes its standby peers list the stream underholderUnresponsiveand reportDegraded(never not ready: a restart cannot free the holder's lock). A holder whose registry writes keep failing looks the same, so a persistently brokennagare_leasestable shows up as degraded standby streams rather than only lost visibility. A holder whose first write failed retries it on every renewal until it lands, taking the row over from the previous holder. During a rolling upgrade from a version before 0.19.0, old holders do not renew, so standby streams on new nodes may reportholderUnresponsiveuntil the old nodes are gone;StandbyHolderStaleAfter = Timeout.InfiniteTimeSpanturns the check off. See stream readiness. - No fencing tokens. This is a visibility layer, not a replacement for distributed locks. The lock primitive itself (Postgres advisory locks, etc.) gives you mutual exclusion;
nagare_leasesonly describes who currently holds it.
When you don't need this
Single-instance deployments don't need the lease registry at all — the dashboard's "this node" data is already enough. The registry exists to answer the question "which of my N nodes is the outbox running on right now?" If N is always 1, leave it off.
If you do turn it on for a single instance, the table is harmless: one or two rows, no contention, sweep does nothing.