Each controller exports metrics and traces through the OpenTelemetry Go SDK, configured by the environment alone; it has no flag of its own for them. A signal is exported only once one of OTEL_<SIGNAL>_EXPORTER, OTEL_EXPORTER_OTLP_<SIGNAL>_ENDPOINT or OTEL_EXPORTER_OTLP_ENDPOINT is set, <SIGNAL> being METRICS or TRACES; with none of them set, or with OTEL_SDK_DISABLED=true, the controller exports nothing and the instruments below are not collected. Once a signal is on, it is sent as OTLP over http/protobuf to https://localhost:4318 unless the variables below say otherwise, and a failed export is logged. controller-runtime’s own Prometheus metrics are served apart, on --metrics-bind-address.
Environment #
| Variable | Default | Effect |
|---|---|---|
OTEL_SDK_DISABLED | false | true, in any case, exports nothing whatever else is set; no exporter is built and no Prometheus listener opened. |
OTEL_SERVICE_NAME | the controller’s name, such as cluster-controller | service.name of every metric and span. |
OTEL_RESOURCE_ATTRIBUTES | unset | Further resource attributes, as key=value pairs separated by commas. |
OTEL_METRICS_EXPORTER | otlp once metrics are on | otlp, prometheus, console or none; set to any but none, it turns metrics on; none exports no metrics whatever else is set. |
OTEL_TRACES_EXPORTER | otlp once traces are on | otlp, console or none; set to any but none, it turns traces on; none exports no spans whatever else is set. |
OTEL_EXPORTER_OTLP_PROTOCOLOTEL_EXPORTER_OTLP_METRICS_PROTOCOLOTEL_EXPORTER_OTLP_TRACES_PROTOCOL | http/protobuf | http/protobuf or grpc. |
OTEL_EXPORTER_OTLP_ENDPOINTOTEL_EXPORTER_OTLP_METRICS_ENDPOINTOTEL_EXPORTER_OTLP_TRACES_ENDPOINT | https://localhost:4318 over http/protobuf, https://localhost:4317 over grpc | Where OTLP is sent; set, it turns on the signals it applies to. An http:// endpoint sends without TLS. |
OTEL_EXPORTER_OTLP_INSECUREOTEL_EXPORTER_OTLP_METRICS_INSECUREOTEL_EXPORTER_OTLP_TRACES_INSECURE | false | true sends OTLP without TLS. |
OTEL_EXPORTER_OTLP_CERTIFICATEOTEL_EXPORTER_OTLP_METRICS_CERTIFICATEOTEL_EXPORTER_OTLP_TRACES_CERTIFICATE | unset | PEM file of the CA certificates the collector’s certificate is verified against. |
OTEL_EXPORTER_OTLP_HEADERSOTEL_EXPORTER_OTLP_METRICS_HEADERSOTEL_EXPORTER_OTLP_TRACES_HEADERS | unset | Headers sent with each export, as key=value pairs separated by commas. |
OTEL_EXPORTER_OTLP_TIMEOUTOTEL_EXPORTER_OTLP_METRICS_TIMEOUTOTEL_EXPORTER_OTLP_TRACES_TIMEOUT | unset | Milliseconds an export may take. |
OTEL_EXPORTER_OTLP_COMPRESSIONOTEL_EXPORTER_OTLP_METRICS_COMPRESSIONOTEL_EXPORTER_OTLP_TRACES_COMPRESSION | none | gzip compresses each export. |
OTEL_EXPORTER_PROMETHEUS_HOSTOTEL_EXPORTER_PROMETHEUS_PORT | localhost, 9464 | Where OTEL_METRICS_EXPORTER=prometheus serves /metrics for scraping. The chart value metrics.prometheus.enabled sets OTEL_METRICS_EXPORTER=prometheus and OTEL_EXPORTER_PROMETHEUS_HOST=0.0.0.0, and puts port 9464 on the metrics Service and ServiceMonitor. |
OTEL_METRIC_EXPORT_INTERVAL | unset | Milliseconds between two exports of the otlp and console metrics exporters, each reading the resources’ status; prometheus reads it at each scrape. |
OTEL_METRIC_EXPORT_TIMEOUT | unset | Milliseconds a metric export may take. |
OTEL_TRACES_SAMPLEROTEL_TRACES_SAMPLER_ARG | parentbased_always_on | Which reconcile spans are kept. |
A variable with METRICS or TRACES in its name applies to that signal alone and overrides the one without.
The OpenTelemetry Operator #
Of the variables above, the chart sets only OTEL_METRICS_EXPORTER and OTEL_EXPORTER_PROMETHEUS_HOST, and only under metrics.prometheus.enabled. The OpenTelemetry Operator sets them from an Instrumentation in the release namespace:
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
name: nats-operator
namespace: nats-operator
spec:
exporter:
endpoint: http://otel-collector.observability.svc:4318
into every pod annotated instrumentation.opentelemetry.io/inject-sdk: "true", which the chart’s podAnnotations carries:
podAnnotations:
instrumentation.opentelemetry.io/inject-sdk: "true"
cluster:
podAnnotations:
resource.opentelemetry.io/service.name: cluster-controller
auth:
podAnnotations:
resource.opentelemetry.io/service.name: auth-controller
jetstream:
podAnnotations:
resource.opentelemetry.io/service.name: jetstream-controller
inject-sdk injects environment only (SDK environment variables only): OTEL_EXPORTER_OTLP_ENDPOINT from spec.exporter.endpoint, which turns metrics and traces on; OTEL_SERVICE_NAME; OTEL_RESOURCE_ATTRIBUTES with the pod’s Kubernetes attributes; OTEL_TRACES_SAMPLER and OTEL_TRACES_SAMPLER_ARG from spec.sampler; and every entry of spec.env, where any other variable above goes, such as OTEL_EXPORTER_OTLP_PROTOCOL: grpc for a collector’s gRPC port. The controllers speak http/protobuf unless told otherwise, so the endpoint is the collector’s OTLP/HTTP port.
OTEL_SERVICE_NAME is the Deployment’s name, <release>-<controller>, unless the pod annotation resource.opentelemetry.io/service.name names another, as above (Configure resource attributes). The three controllers’ pods share their app.kubernetes.io/name and app.kubernetes.io/instance labels, so spec.defaults.useLabelsForResourceAttributes alone gives all three one service.name.
To send through a collector in each controller’s pod instead, create an OpenTelemetryCollector (opentelemetry.io/v1beta1) with mode: sidecar in the release namespace, whose otlp receiver takes HTTP on port 4318; add sidecar.opentelemetry.io/inject: "true" to podAnnotations; and set the Instrumentation’s spec.exporter.endpoint to http://localhost:4318 (Sidecar injection).
Both annotations are read off the pod: on the Deployment, under the chart’s annotations, the OpenTelemetry Operator does not see them. The fields and annotations here are those of OpenTelemetry Operator v0.159.0 (Instrumentation API reference).
Metrics #
The gauges are read off the resources’ status at each export. Every point carries kind, namespace and name of the resource it describes. The Prometheus column is each instrument’s name as OTEL_METRICS_EXPORTER=prometheus serves it on port 9464, which metrics.prometheus.enabled sets.
| Instrument | Prometheus | Type | Unit | Controllers | Attributes | Reads | Value |
|---|---|---|---|---|---|---|---|
nats_operator.condition | nats_operator_condition | gauge | cluster-controller, auth-controller, jetstream-controller | kind, namespace, name, type, reason | status.conditions of every kind the controller reconciles | 1 while the condition is True, 0 while it is False or Unknown. | |
nats_operator.account.jwt_expiry | nats_operator_account_jwt_expiry_seconds | gauge | s | auth-controller | kind, namespace, name | NatsAccount status.jwt | When the NatsAccount’s current JWT expires, in seconds since the Unix epoch; a JWT that never expires has no point. |
nats_operator.rollout.pending_servers | nats_operator_rollout_pending_servers | gauge | {server} | cluster-controller | kind, namespace, name | NatsCluster status.rollout.pending | Servers still to update in the NatsCluster’s rollout. |
nats_operator.rollout.gate | nats_operator_rollout_gate | gauge | cluster-controller | kind, namespace, name, waiting_for | NatsCluster status.rollout.gate.waitingFor | 1 while the rollout’s gate is closed, by what it waits for. | |
nats_operator.balancer.leader_skew | nats_operator_balancer_leader_skew | gauge | {leader} | jetstream-controller | kind, namespace, name, pool | NatsBalancer status.pools[].leaderSkew, NatsSystemBalancer status.skew.leaders | The most leaders one server carries less the fewest another does: per pool for a NatsBalancer, over the NATS cluster for a NatsSystemBalancer. |
nats_operator.balancer.pending_moves | nats_operator_balancer_pending_moves | gauge | {move} | jetstream-controller | kind, namespace, name, move_kind | NatsSystemBalancer status.pending | Moves the NatsSystemBalancer requested that are not yet complete, by kind of move. |
nats_operator.balancer.held_passes | nats_operator_balancer_held_passes_total | counter | {pass} | jetstream-controller | kind, namespace, name, reason | NatsBalancer and NatsSystemBalancer status.conditions[Holding], after each pass | Balancer passes that ended with Holding True, by its reason. |
nats_operator.evacuation.remaining | nats_operator_evacuation_remaining | gauge | {stream} | jetstream-controller | kind, namespace, name | NatsClusterEvacuation status.remaining | Streams still to leave the evacuation’s source NATS cluster. |
nats_operator.evacuation.stale_placements | nats_operator_evacuation_stale_placements | gauge | {stream} | jetstream-controller | kind, namespace, name | NatsClusterEvacuation status.stalePlacement | Moved streams no resource owns whose config still names the source NATS cluster. |
Account JWT expiry #
The auth controller re-signs each account JWT at half its jwtTTL, so while it runs, no account JWT expires sooner than half the shortest jwtTTL from now: 24h under the default 48h. Scraped through the chart’s ServiceMonitor, this fires once one expires within 23h:
min(nats_operator_account_jwt_expiry_seconds) - time() < 23 * 3600
The threshold assumes the default 48h jwtTTL; a shorter jwtTTL needs one below half of it.
The gauge is exported by the auth controller itself and goes stale once its scrape fails, so that expression returns nothing while the auth controller is down. This fires then, from the ServiceMonitor’s endpoint otel-metrics:
absent(up{job=~".+-auth-controller-metrics", endpoint="otel-metrics"} == 1)
Traces #
Each reconcile runs in a span named Reconcile <kind>, such as Reconcile NatsCluster, carrying kind, namespace and name. A reconcile that returns an error marks its span failed with it.
Events #
The controllers record these through the events.k8s.io API, regarding the resource named.
| Reason | Type | Controller | Regarding | Recorded when |
|---|---|---|---|---|
MoveStarted | Normal | jetstream-controller | NatsBalancer, NatsSystemBalancer, NatsClusterEvacuation | A balancer moves a stream’s leader or placement, or an evacuation requests a stream’s move. |
MoveDone | Normal | jetstream-controller | NatsSystemBalancer, NatsClusterEvacuation | A move the system balancer or the evacuation requested is seen complete. |
MoveCancelled | Normal | jetstream-controller | NatsClusterEvacuation | Deleting an evacuation cancels a move still in flight. |
MoveRefused | Warning | jetstream-controller | NatsClusterEvacuation | The server refuses a move the evacuation requests. |
EvacuationRefused | Warning | jetstream-controller | NatsClusterEvacuation | The evacuation refuses to start because a server of its source NATS cluster carries the target tags. |
RolloutStep | Normal | cluster-controller | NatsCluster | A rollout restarts a server, or starts removing or replacing one. |
GateBlocked | Warning | cluster-controller | NatsCluster | A rollout’s gate has been closed long enough for Progressing to read GateBlocked. |
ReconcileFailed | Warning | cluster-controller | NatsCluster | A reconcile fails on anything but a write conflict; Progressing reads ReconcileFailed with the same message. |
JWTPushed | Normal | auth-controller | NatsSystemAccount, NatsAccount | A newly signed account or system account JWT is pushed to the servers. |
JWTHeld | Warning | auth-controller | NatsOperator, NatsAccount | An account JWT is not signed because its revocations cannot be recovered from the servers. |
UserKicked | Normal | auth-controller | NatsUser | A deleted user’s live connections are closed. |