Three Kubernetes controllers that between them deploy NATS clusters, own their auth plane, and manage and balance JetStream. Vocabulary is CONTEXT.md’s; “operator” unqualified is not used. Every nats-server citation is against v2.15.0 unless stated.
1. Scope #
In scope: NATS clusters and superclusters spanning Kubernetes clusters, leaf nodes, the auth plane (NATS operator, system account, accounts, users), JetStream resources (streams, consumers, key-value buckets, object stores), balancing, evacuation, rollout, scale-down and server replacement.
Out of scope:
- Spreading client connections after a rollout.
- Delivering creds past the Kubernetes cluster boundary; External Secrets, SOPS or similar carry them.
- Node pools, storage classes, DNS zones; GitOps around the controllers provides them.
- General self-healing beyond what rollout, scale-down and balancing need.
- Moving a stream to a NATS cluster outside its supercluster.
- Migrating any existing deployment onto these controllers.
- A KMS integration for signing keys.
- Rotating identity keys: a new identity is a new NATS operator or account, which is a migration.
2. Architecture #
2.1 Controllers and API groups #
Three separately deployable controllers, each owning one API group, plus one shared group every controller reads. All groups are v1beta1, every kind is namespaced, and every kind name is prefixed Nats. An API group marks one controller’s boundary for installation, RBAC and versioning, so a group with one kind is not a smell (ADR 0001).
| Group | Owner | Kinds |
|---|---|---|
nats.mikluko.io | shared | NatsReferenceGrant, NatsOperatorTrust, NatsAccountTrust, NatsConnection |
cluster.nats.mikluko.io | cluster controller | NatsCluster |
auth.nats.mikluko.io | auth controller | NatsOperator, NatsSystemAccount, NatsAccount, NatsUser |
jetstream.nats.mikluko.io | JetStream controller | NatsStream, NatsConsumer, NatsKeyValue, NatsObjectStore, NatsBalancer, NatsSystemBalancer, NatsClusterEvacuation |
2.2 What each controller reads #
No controller reads another controller’s group. Every kind two controllers share lives in nats. (ADR 0002).
- Cluster controller:
cluster.*, andNatsOperatorTrust,NatsAccountTrust,NatsConnection,NatsReferenceGrant. - Auth controller:
auth.*andnats.*; it writes the status of reference-form trust objects (section 5.1). - JetStream controller:
jetstream.*, andNatsConnection,NatsReferenceGrant.
The JetStream and auth controllers act on NATS clusters the cluster controller did not deploy as readily as on managed ones.
2.3 What each kind does #
| Role | Kinds | What exists because of it |
|---|---|---|
| Creates | NatsCluster | one StatefulSet and ConfigMap per server, Services, PDB, self-signed route certificates |
| Creates | NatsOperator, NatsSystemAccount, NatsAccount, NatsUser | keys (unless supplied), signed JWTs pushed to resolvers, creds Secrets |
| Creates | NatsStream, NatsConsumer, NatsKeyValue, NatsObjectStore | objects on a NATS server |
| Acts | NatsSystemBalancer, NatsBalancer, NatsClusterEvacuation | nothing lasting; leader and placement moves |
| Data | NatsConnection, NatsReferenceGrant, NatsOperatorTrust, NatsAccountTrust | nothing; read by the controllers |
There is no supercluster kind: each NatsCluster carries its gateway list, and the supercluster is what those lists form (ADR 0006).
3. Resource API #
The API is specified by the user stories under docs/content/docs/stories/, whose manifests are the source of truth; this section states only the rules the manifests do not show on their own.
| Story | Covers |
|---|---|
| 01-quickstart | NatsCluster, NatsConnection, NatsStream; rollout status |
| 02-auth-plane | NatsOperator, NatsSystemAccount, NatsAccount, NatsUser, NatsOperatorTrust; bring-your-own-key users |
| 03-unmanaged | JetStream against a NATS cluster nobody here deployed; adoption; NatsConsumer, NatsKeyValue, NatsObjectStore |
| 04-team-self-service | namespaces and NatsReferenceGrant |
| 05-account-wiring | exports and imports |
| 06-supercluster | gateways, literal NatsOperatorTrust, remote controllers’ users |
| 07-balancing | NatsSystemBalancer, NatsBalancer, pools, the stepdown export preset |
| 08-stream-transfer | moving a stream between NATS clusters of a supercluster |
| 09-acceptance | a three-cluster production supercluster |
| 10-leafnodes | leaves, with and without the hub’s auth plane; NatsAccountTrust |
| 11-evacuation | NatsClusterEvacuation; the NatsCluster deletion guard |
Rules across kinds:
- A Secret is always read through
secretKeyRef: {name, key}.keydefaults touser.credsfor credentials andca.crtfor a CA. - A
NatsUserwrites creds, optionally, in the shape aNatsConnectionreads:credentials.secretKeyRef. - A reference names its kind wherever more than one kind can answer it.
- Mutually exclusive fields (
publicKeyandcredentials,accountRefandpublicKey,streamandstreamRef,keys.identityandjwt,presetandpermissions) are refused together at apply by CEL. - A one-shot imperative is an annotation the controller clears once it has acted (
force-step,replace-server); desired state is always a typed field.force-deleteis read only while itsNatsClusteris being deleted, and never cleared. - TLS is optional on every listener (client, routes, gateways, leafnodes). Route TLS is on by default, self-signed when no certificate is named, and can be turned off.
4. Cluster controller #
4.1 Workload #
Each server is its own StatefulSet with replicas: 1, its own rendered ConfigMap, and an explicit server_name (ADR 0004). The controller also owns the client and headless Services, a PodDisruptionBudget with maxUnavailable: 1 across the NATS cluster’s pods, and the external gateway and leafnode Services rendered from the templates in spec. It writes no object of those names that exists and it does not control; the NatsCluster then reads Ready=False, reason: ReconcileFailed, naming every such object. The servers’ pods prefer distinct nodes unless podTemplate sets affinity. Pod and PVC templates pass through; GOMEMLIMIT and the JetStream memory store derive from resources.limits.memory and the file store from the PVC size, each overridable under jetstream.limits. A prometheus-nats-exporter sidecar, its image pinned by digest and its resources small unless exporter.resources sets them, is on by default; exporter.enabled: false turns it off. tls puts the client listener under TLS, its certificate from secretRef or from certManager issued for the client Service’s names; clients then dial tls://, as status.endpoints.client reads, and need the CA that issued it, and the controller’s own system connection trusts the Secret’s ca.crt, or the system roots without one. Unset, clients connect in the clear. The monitoring port, which has no authentication, is on the headless Service and never the client Service; it stays reachable at each pod’s address, so the controller renders a NetworkPolicy over the pods that admits it only from the cluster controller’s own namespace. The same policy admits the exporter’s metrics port, which republishes what the monitoring port serves, from that namespace and the peers exporter.from names; the route port, which has no authentication of its own, only from the NATS cluster’s own pods; and every other port the controller renders from anywhere. The exporter reads it over the pod’s loopback, which no NetworkPolicy governs. monitor.networkPolicy: false renders none; a port podTemplate adds is not admitted by it. A self-signed route certificate is valid for a year and reissued under the same CA key once a third of that remains, the key kept in the Secret <name>-routes-ca, which no pod mounts. Pods meet the restricted Pod Security Standard and mount no ServiceAccount token, but podTemplate merges over that too, so whoever may write a NatsCluster runs pods with any privilege its namespace admits.
spec.version names the nats-server version the controller renders for; image sets the repository and a digest, never the tag. 2.15.0 is the minimum: lower is refused at apply. Any patch change is allowed, and so is one minor up or down; a move of two or more minors, either way, is refused at apply. The hazard is the Raft log, not the file store: a newer minor writes log entries an older one may not replay (2.15’s deleteRangeOp panics 2.12), and upstream reads a new entry one minor before it writes it. Nothing constrains versions across a supercluster. A config key the version does not know fails the server’s config load rather than warning (opts.go:1920).
4.2 Config changes #
A rendered-config change reloads a server where every changed field is in the reload allow-list for its version, and restarts it otherwise. A reload succeeds or fails whole: one restart-only field rejects the batch. The allow-list mirrors diffOptions per supported version. Restart-only fields include gateway remotes, route and gateway listen addresses, JetStream store_dir, lame-duck settings, system_account (reload.go:1969-1973), the trusted NATS operators (reload.go:1619), an account preloaded from a NatsAccountTrust, and any resolver change: a reload accepts one but installs a resolver it never starts (reload.go:2236), so a claims update still lands in the old directory while lookups read the new one. Routes, TLS certificate and key, server_tags, authorization, max_payload, and raised JetStream limits reload. So does the preloaded system account JWT, which a revocation or a jetstream-stepdown export re-signs: a reload keeps the preload it started with (reload.go:1614, the preloads being unexported), the re-signed JWT reaches running servers as a claims update, and the preload is read again only at the next start. Leafnode remotes can be added and removed by reload (reload.go:931-990).
The controller requests a reload with $SYS.REQ.SERVER.<id>.RELOAD over its system user, confirms it by /varz config_digest and config_load_time, and falls back to a restart on any error.
4.3 Trust, resolver, gateways, leaf nodes #
Every NatsCluster with an auth plane names its trust roots with auth.trustRef to a NatsOperatorTrust (section 5.1), and its own system credentials with auth.systemCredentials. It renders operator, system_account, and a full resolver with the system account’s JWT preloaded, so a server boots before its gateways are up. auth.resolver is Full or Cache, Full by default except on a leaf that preloads no account; Memory is not offered because accounts arrive over the network.
Gateways: gateway.remotes lists every member of the supercluster (a NATS cluster’s own entry is skipped), gateway.discovery is Explicit (every remote, reject_unknown on) or Gossip (seeds, reject_unknown off; gossip discovery then completes the mesh). The same list is in every member, kept alike by GitOps. Any change restarts servers: gateway config is not reloadable (reload.go:1765-1790). gateway.service renders the external Service; gateway.advertise is the address the server advertises. Gateway TLS, from secretRef or certManager, verifies peers both ways against the Secret’s ca.crt, and a Secret without one holds the servers with Progressing=True, reason: GatewayCertificateNotReady. A gateway without tls is refused, Ready and Progressing False with reason GatewayWithoutTLS and nothing rendered, unless the cluster controller runs with --allow-gateway-without-tls. nats-server demands a certificate from every inbound gateway (opts.go:3310), and without a CA file verifies it against the system roots, so a certificate from any public issuer would join the supercluster, which has no other authentication; a certificate from ACME, which issues no ca.crt, is therefore not accepted.
Leaf nodes, hub side: leafnodes with tls, service and advertise, refused without auth (CEL-enforced), since a listener without an auth plane admits any leaf into the global account. Leaf side: leafRemotes[], each naming a NatsConnection (URL, CA, credentials) and optionally localAccountTrustRef or localSystemAccount: true; no localAccount is offered, since a leaf without an auth plane has only the global account. The remotes’ creds and CAs go into a Secret projected into the config volume beside nats.conf, so kubelet refreshes both together and a remote added or removed reloads, where auth.systemCredentials gives the controller a system user to request the reload with; without one every config change restarts. A remote whose NatsConnection no grant admits is rendered out of the config and the Secret, the other remotes kept, and LeafnodesConnected reads False with the refusal naming it. A leaf running JetStream must set jetstream.domain (CEL-enforced). A leaf that trusts the hub’s NATS operator needs a remote bound to its system account, authenticated as a user of the hub’s system account; without that remote it cannot fetch an account it has not cached (auth.go:1020, leafnode.go:721). Its resolver is Cache when it preloads no account. A leaf preloading an account through a literal NatsAccountTrust runs Full on jetstream.volumeClaimTemplate and is refused without one unless it sets Cache: a fetch never replaces a preload, only Full’s sync with the hub or a claims push does, and a persistent directory keeps what was synced across a restart. The preload is a copy, valid only within the account’s jwtTTL unless that account sets jwtTTL: 0. LeafnodesConnected and status.leafRemotes count the servers holding each remote, from LEAFZ over the system user or /leafz; remotes binding the same local account are told apart by count alone, since LEAFZ names no remote.
4.4 Rollout #
One server at a time: a step writes the server’s ConfigMap and its pod template, which names the revision, and its StatefulSet restarts the pod. Order: highest ordinal first, the meta leader’s server last. The controller moves no leader before a restart: the preStop enters lame-duck mode (nats-server --signal ldm=<pid file>), which steps down every Raft leader on the server and makes each of its nodes an observer before any client is told (server.go:4472, raft.go:913-936). The rendered lame_duck_duration of two minutes ends inside the pod’s terminationGracePeriodSeconds of 300.
The gate to the next server: every server’s pod is Ready, every server on the target revision reports it, with JetStream every server that answers is a member of the meta group, and the NATS cluster is Settled. A meta leader in another NATS cluster of the supercluster is asked for its peers, and the meta group’s members are this NATS cluster’s servers among them; when that leader does not answer, the meta group is judged from its followers (FromFollowers), and every server that answers counts as a member. A closed gate reads Progressing=True, reason: RollingRestart, and after ten minutes GateBlocked, naming the groups that are not current, while status.rollout.gate.since says since when; it holds indefinitely, and nothing proceeds on a timeout. spec.rollout.paused stops a rollout before its next step; the annotation cluster.nats.mikluko.io/force-step: "<server>" restarts that waiting server now, through a closed gate and through paused, and is cleared once read. The same gate holds the next voluntary step after any involuntary disruption.
4.5 Scale-down and server replacement #
Scale-down and replacement are steps of the rollout (section 4.4), under its gate, one server at a time: servers beyond replicas first, highest ordinal first, then restarts, then replacements, the meta leader’s server last. A removal once begun is carried through, paused or not.
Scale-down removes the servers beyond replicas: $JS.API.SERVER.EVACUATE one, wait for it to hold no Raft group and for Settled apart from itself, step it down if it leads the meta group, $JS.API.SERVER.REMOVE it, then delete its StatefulSet, its PVC only when the PVC carries the server’s labels, and its ConfigMap only when the NatsCluster controls it; the NatsCluster then reads Ready=False, reason: ReconcileFailed, naming each one left. A stream whose replica count exceeds the new size stops the shrink with Progressing=False, reason: ScaleDownBlocked; nothing is forced. Without JetStream a server is deleted with no evacuation; with JetStream and no auth.systemCredentials nothing can be evacuated, and scale-down and replacement are blocked.
A change to jetstream.volumeClaimTemplate (class, size, zone) replaces servers: evacuate, remove, delete that server’s StatefulSet and, under the same rule, its PVC, and recreate it under the same server_name once the old PVC is gone. The annotation cluster.nats.mikluko.io/replace-server: "<server>" replaces one server on demand. A server’s peer ID is a hash of its name (server.go:4079-4081), so it rejoins as the same peer, but only after the five-minute removal tombstone lapses (raft.go:316, readmission raft.go:4121-4169, measured at 5m1s). The fresh pod starts at once and answers outside the meta group; the rollout gate holds on it until the meta leader lists it again. Progressing reads ScalingDown or ReplacingServer while a step runs.
4.6 Deletion guard #
Deleting a NatsCluster with JetStream that still holds stream groups waits, with Deleting=True, reason: JetStreamDataRemains naming them, found through the controller’s own system JSZ; a NATS cluster that cannot be observed waits too, with reason: ObservationFailed. The annotation cluster.nats.mikluko.io/force-delete overrides both. NatsClusterEvacuation (section 6.6) is how the data leaves first.
4.7 Status #
Ready, Settled, Progressing, and where they apply GatewaysConnected and LeafnodesConnected; endpoints (client, monitor, gateway), the config revision and how it was applied, the rollout block (updated, current, pending servers, the gate and since when), and one entry per server.
5. Auth controller #
The auth controller runs in the home cluster only, and stops at JWT claims: it configures nothing about JetStream beyond the limits signed into account JWTs.
5.1 Keys and trust #
Keys (ADR 0005). NatsOperator, NatsSystemAccount and NatsAccount take optional keys: {identity: {secretKeyRef}, signing: [{name, secretKeyRef, retiring}]}. Set, the keys are adopted; an adopted account keeps its public key, so user JWTs already issued under it keep working without NatsUser resources, while its claims are re-signed from spec. Omitted, the controller generates keys into Secrets named <name>-<operator|systemaccount|account>-<identity|signing-1> and annotated auth.nats.mikluko.io/generated-for: <operator|systemaccount|account>/<name>, with no owner reference, so they outlive the object and an object of the same kind and name applied again takes the same keys; it reads a Secret under such a name only where that annotation names the object; any other is Ready False, reason SecretConflict.
Offline identities. NatsOperator.spec.jwt holds a NATS operator JWT signed elsewhere, exclusive with keys.identity; an account’s publicKey is exclusive with keys.identity. Either requires at least one signing key. A NatsAccount no grant admits to its NatsOperator is refused before its key is recorded, Ready False, reason ReferenceNotPermitted. Deleting that grant once the account is signed deletes it from the servers through its NatsOperator’s status.deletedAccounts and empties its status.jwt once that list records it. A NatsUser of an account not admitted is not signed, Ready False, reason AccountNotAdmitted. An account’s public key is unique under its NatsOperator: a NatsAccount whose key is the identity of the NatsSystemAccount that NatsOperator references, as its status, publicKey or identity seed Secret gives it, or another NatsAccount under it records, is not signed, Ready False, reason PublicKeyInUse, naming the holder; where two accounts record one key, the older keeps it. nats-server takes any account JWT its NATS operator signed for the system account’s key in place of the system account’s own and cuts off the system account’s users (pinned in-process by TestResolvers_SystemAccountKeyPushed), so this refusal is what keeps a namespace granted the NatsOperator out of the system account. The Kubernetes cluster then holds signing seeds only: one NATS operator signing key for account JWTs (the system account’s included), and one per account for its users and its activation tokens, which an account signing key may sign (activation_claims.go:34 in jwt/v2, accepted at accounts.go:3094-3108).
Custody. Seeds live only in Kubernetes Secrets, never in spec or status. Through secretKeyRef those Secrets may come from an external store, and that copy is the backup; losing the home namespace without one loses the ability to change any account. A generated identity seed is generated once: where its Secret is gone while the object’s status.publicKey records the identity, the object reads Ready False, reason SeedLost, and no new identity replaces it, since one would orphan every JWT issued under the old.
Rotation. A signing key is added, the controller re-signs what it manages (users, or for the NATS operator, accounts) and rewrites creds Secrets, then the key marked retiring: true is removed, which invalidates everything still signed by it. Any change to the NATS operator JWT (signing keys, system account) is restart-only (reload.go:1619): the NatsOperatorTrust copies are updated first, then every server in the supercluster rolls. With an offline identity each such change also needs the JWT re-signed offline.
Trust objects (ADR 0003). NatsOperatorTrust holds either operatorRef (same Kubernetes cluster; the auth controller writes the NATS operator and system account JWTs into its status) or the literal operatorJWT and systemAccountJWT, copied by GitOps wherever no auth controller runs. NatsAccountTrust likewise holds accountRef or a literal publicKey and optional jwt, which a leaf preloads. NatsOperator.spec.systemAccountRef selects the live system account among any number of NatsSystemAccounts; an unreferenced one is not signed.
5.2 Distribution #
Account JWTs reach servers through the full resolver: $SYS.REQ.CLAIMS.UPDATE crosses gateways (every account runs interest-only over gateways since 2.9.0) and restarts nothing. The system account JWT is signed into the NatsOperator status and pushed as the NatsSystemAccount’s own, whose status.distribution and JWTPushed event report it; it is not pushed while that NatsOperator reads RevocationsUnrecovered True, and is signed again where a server holds a newer one. status.distribution reports how many servers hold the current JWT, built from the STATSZ roster and each server’s CLAIMS.LOOKUP, since the server offers no aggregate. Without --system-connection nothing is pushed and no deleted user is kicked, and every account and user says so with Distributed False, reason NoSystemConnection.
Deletes. The resolver’s catch-up sync only adds, so an account deleted while a server was away survives there. Two measures close it: the controller re-sends $SYS.REQ.CLAIMS.DELETE for deleted accounts whenever a server or NATS cluster rejoins (seen in STATSZ), and every account JWT carries an expiry, NatsAccount.spec.jwtTTL (default 48h), re-signed at half its TTL. A home outage therefore costs nothing for half a TTL and expires accounts after a whole one. System account JWTs carry no expiry, because their literal copies in NatsOperatorTrust would otherwise go stale; for the same reason jwtTTL: 0 signs an account JWT with no expiry, which an account preloaded by a leaf through a literal NatsAccountTrust sets, since that copy otherwise stops working a TTL after it was taken. The controller pushes CLAIMS.UPDATE for every change and never an older JWT: a push carries no iat guard, and an older JWT pushed over a newer one stays.
5.3 Users #
A NatsUser takes permissions or a preset: cluster-controller, jetstream-controller and the auth controller’s own (claims updates, the CONNZ ping, KICK) for system users; readonly, leafnode for ordinary accounts. A controller preset grants exactly the subjects its controller requests, listed on the NATS permissions page, and subscribes only under its own inbox, _INBOX.<preset>.>, which its controller dials with, so no controller reads another’s replies. It writes creds to credentials.secretKeyRef if given, or, with publicKey instead, receives only a signed JWT in status and no Secret. A user’s public key is unique in its account: a publicKey another NatsUser of the account records is not signed, Ready False, reason PublicKeyInUse, naming the holder, and deleting that user revokes nothing; where two record one key, the older keeps it. A grant admitting NatsUsers to an account lets the grantee claim and revoke any user key of that account, keys issued outside the controller included, since no NATS record tells those apart. Controllers outside the home cluster run as system users declared at home, scoped by preset, never expiring, and carried across by External Secrets or SOPS.
Deletion: the user’s public key is added to the account JWT’s revocations and pushed; a finalizer holds the resource until distribution shows every server current; the controller then finds the user’s live connections with a system $SYS.REQ.SERVER.PING.CONNZ (events.go:1286), which each server filters by account and user public key (monitor.go:407-413) and kicks each with $SYS.REQ.SERVER.<id>.KICK (events.go:62); the creds Secret goes last. Revocations are recorded in the NatsAccount or NatsSystemAccount status.revocations, each with the signing keys that may have issued a revoked JWT, and the account JWT is signed from that record and the JWT it replaces, so losing either keeps them. Losing both, with the whole status, the controller takes them back from the JWT the servers hold ($SYS.REQ.ACCOUNT.<key>.CLAIMS.LOOKUP, accounts.go:4579) before signing again. While any server of the STATSZ roster does not answer, since it may hold a newer JWT than those that do, an account whose status.distribution records it distributed is not signed, Ready False with reason RecoveringRevocations; any other, a new account among them, is signed with the condition RevocationsUnrecovered True, and the servers are asked again once they answer. A revocation is dropped once none of those signing keys is left on the account, or once the revoked JWT has expired, which for users is never unless a user expiry is added; rotating an account’s signing key is what resets its list. A user’s accountRef cannot change. A user whose key changes, by a new publicKey or a creds Secret that had to be written afresh, has the key it held revoked from that moment through the same record, and the server closes that key’s connections when it takes the account JWT (accounts.go:3994-4026); status.replacedKeys holds each such key until the account JWT revokes it, and a deleted user’s finalizer waits on those too. publicKey stays mutable, since a client rotating its own key is ordinary and a creds Secret can be rewritten under any rule on spec.
5.4 Exports and imports #
An import names an export, not a subject; the subject, type and response type come from the exporting account. A private export lists its importers, and the controller mints each an activation token. A cross-namespace import needs both a NatsReferenceGrant in the exporter’s namespace and, for a private export, the importer listed. The export preset jetstream-stepdown expands to service exports of $JS.API.STREAM.LEADER.STEPDOWN.* and $JS.API.CONSUMER.LEADER.STEPDOWN.*.*, and the controller adds the matching prefixed imports to the system account’s JWT; that is what lets the system balancer move the account’s leaders, in a supercluster as in a lone NATS cluster.
6. JetStream controller #
6.1 One mode #
The JetStream controller reaches every NATS cluster, managed or not, through a NatsConnection: servers, tls.ca, credentials. The credentials decide the account, so every JetStream resource names only connectionRef (ADR 0002). The system account cannot create, update, delete or step down another account’s streams: the whole $JS.API.> surface resolves the caller’s account (jetstream_api.go:1107, :955-969), so a per-account user is required for everything but observation, placement moves and evacuation.
6.2 Kinds #
NatsStream, NatsConsumer, NatsKeyValue, NatsObjectStore. Each kind’s fields are the corresponding NATS config in camelCase, beside connectionRef and the lifecycle policies of section 6.3: StreamConfig and ConsumerConfig from nats-server, KeyValueConfig and ObjectStoreConfig from nats.go v1.54.0 (jetstream/kv.go, jetstream/object.go), whose bucket is name here. A user-supplied metadata map is merged with the ownership marker, which the controller owns. Mirrors and sources are stream and key-value fields. One consumer kind: push when deliverSubject is set, pull otherwise. A consumer names its stream by stream (server-side name, for a stream with no resource) or streamRef (it then waits for the stream and inherits its connection). connectionRef, and a consumer’s stream and streamRef, cannot change. A server-side name defaults to metadata.name.
6.3 Lifecycle policies #
Typed fields, beside connectionRef:
adoptionPolicy: Never | Adopt | AdoptOrCreate(defaultNever), with ACK’s semantics:Adoptrequires the object to exist and writes its config into spec;AdoptOrCreateapplies the spec and late-initializes omitted fields;Nevergoes Terminal on an object it does not own.deletionPolicy: Retain | Delete, defaultRetainfor streams, key-value buckets and object stores,Deletefor consumers.terminalPolicy: Hold | Retry(defaultHold): whether a Terminal condition waits for an edit or is rechecked every resync.
A resource being deleted whose NatsConnection no longer exists, or is no longer admitted by a grant, is released without its deletion policy running, and its server object stays.
Ownership is a marker naming the resource’s UID in the object’s own metadata map (stream.go:129, consumer.go:133).
6.4 Drift and immutable fields #
The controller re-reads on a resync period and corrects drift by reapplying spec, reporting Synced=False, reason: DriftCorrected once. Fields nats-server will not change are refused at apply by CEL transition rules: for streams name, storage, retention to or from WorkQueue, mirror, allowMsgCounter, persistMode, and the one-way sealed, denyDelete, denyPurge, allowMsgTTL, allowMsgSchedules (stream.go:2542-2604); for key-value buckets and object stores storage, since each is a stream; for consumers deliverPolicy, ackPolicy, replayPolicy, start sequence and time, heartbeats, flow control, maxWaiting (consumer.go:2542-2594). NatsConsumer alone offers recreateOnImmutableChange.
6.5 Balancing #
Two layered balancers. NatsSystemBalancer (at most one per NATS cluster, system credentials) evens leaders and copies over every account: placement moves by $JS.API.ACCOUNT.STREAM.MOVE.<account>.<stream>, which accepts the system account for any account (jetstream_api.go:3188) and picks servers matching the stream’s configured placement tags as well as the request’s (jetstream_api.go:3269-3280), and leader moves through each account’s jetstream-stepdown export; status.capabilities reports leader: Partial for accounts without it. NatsBalancer (per account, its own connection) evens within pools and yields to the system balancer, including any stream the system balancer has a move pending on.
A pool selects NatsStream, NatsKeyValue and NatsObjectStore resources by label; a stream matching several pools belongs to the first and status reports Overlapping; streams in no pool, and streams with no resource, form the default pool.
A balancer’s scope and boundary is the NATS cluster its connection reaches: it judges only groups placed there, its Settled gate reads only that NATS cluster (a cluster-filtered PING.JSZ), balancers in different Kubernetes clusters do not coordinate, and none moves the meta group. A declared placement.cluster pins a stream; balancers never move one across NATS clusters, and never widen or override declared placement. Leader moves are on by default; placement moves are opt-in. One move per pass, none while unsettled; interval paces it. One balancer at a time moves on a NATS cluster: each takes the NATS cluster’s move lease, kept in the memory of the leader-elected JetStream controller, before a move and holds it until a pass of its own finds the move done, while the others hold with Holding reason MoveLeaseHeld; a lease lapses a minute after its holder’s last pass. During a rollout a balancer pauses per step through its own Settled gate.
6.6 Evacuation #
NatsClusterEvacuation (system credentials, whole NATS cluster only; its connectionRef, from and to cannot change) moves every stream, key-value bucket and object store in every account off from.cluster, save those whose resource sets placement.cluster, to servers matching to.serverTags, one ACCOUNT.STREAM.MOVE per stream with those tags as ephemeral placement; peer selection falls back to any NATS cluster matching them (jetstream_api.go:3280-3313), and consumers follow (jetstream_cluster.go:9859). SERVER.EVACUATE and STREAM.PEER.EVACUATE cannot do this: they stay in the stream’s NATS cluster (jetstream_cluster.go:9723-9763). It refuses to start if a source server carries the target tags. While a server of the source does not answer, the meta group has no leader, or its leader reports a server of the NATS system offline, it makes no move, counts none done and holds with reason ServersDown. A resource whose own spec sets placement.cluster is never moved: one naming the source is listed in status.pinned, and the evacuation never becomes Ready while any remains; the rest are left to their owners. The tags are not written to a moved stream’s config, and a config that still names the source is pulled back by its owner’s next placement edit while the source exists; the evacuation rewrites nothing and lists moved streams with no resource whose config names the source in status.stalePlacement. Deleting it mid-run cancels pending moves with CANCEL_MOVE. Balancers skip streams under evacuation.
7. Tenancy #
Every kind is namespaced. A reference crosses namespaces only where a NatsReferenceGrant in the target namespace admits it (from: group, kind, namespace; to: group, kind, optional name), checked on every reconcile, so deleting a grant revokes what it admitted (ADR 0007). Withdrawing a grant that admits a NatsCluster to a NatsOperatorTrust or a NatsAccountTrust holds the NatsCluster at its last render instead, since both carry only public material. There is no admission webhook. Creds Secrets land in the user’s namespace. NatsSystemAccount is a kind of its own so RBAC can grant it to the platform team alone. A grant admitting a NatsCluster’s leaf remote to a NatsConnection hands that connection’s credentials to the NatsCluster’s namespace, where they are copied into the Secret <name>-leaf-remotes. Whichever NatsAccount records an account key first holds it, so a grant admitting NatsAccounts to a NatsOperator lets the grantee take any account key no NatsAccount records yet, an adopted account’s before its own NatsAccount reconciles included. A preset scopes subjects, not NATS clusters: $JS.API.SERVER.REMOVE names its peer in the payload, not the subject, so a cluster-controller user in one Kubernetes cluster can remove a server from any NATS cluster of the supercluster.
8. Guarantees #
The design promises about client behaviour during leader moves, placement moves, rollouts and evacuations exactly what NATS promises, and no more. A prototype at 200 msg/s on 2.15.0 saw no acknowledged message lost or delivered twice, at most one publish-ack timeout per move absorbed by a Nats-Msg-Id retry, and one Leadership Changed on pull consumers per leader change; that is evidence, not contract.
9. Observability #
Every resource reports conditions. The controllers’ own metrics carry skew, pending moves and held passes.