A platform engineer has Prometheus, run by prometheus-operator in monitoring, scrape the controllers’ metrics, and wants every scrape to verify the certificate it is served. The controllers are to connect to nothing but the API server, DNS and the NATS clusters in nats-system.
The certificate #
Each controller serves its metrics under the certificate in one Secret in the release namespace, which must name every enabled controller’s metrics Service. For the release nats-operator in nats-operator, a Certificate from ca-issuer, an existing cert-manager CA ClusterIssuer, which writes its CA to the Secret’s ca.crt:
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: nats-operator-metrics
namespace: nats-operator
spec:
secretName: nats-operator-metrics-tls
issuerRef:
kind: ClusterIssuer
name: ca-issuer
dnsNames:
- nats-operator-cluster-controller-metrics.nats-operator.svc
- nats-operator-auth-controller-metrics.nats-operator.svc
- nats-operator-jetstream-controller-metrics.nats-operator.svc
The chart #
metrics:
service:
enabled: true
tls:
secretName: nats-operator-metrics-tls
scraper:
serviceAccount: monitoring/prometheus
serviceMonitor:
enabled: true
networkPolicy:
enabled: true
from:
- namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: monitoring}}
egress:
- to: [{ipBlock: {cidr: 10.96.0.1/32}}]
ports: [{protocol: TCP, port: 443}]
- to: [{ipBlock: {cidr: 172.18.0.2/32}}]
ports: [{protocol: TCP, port: 6443}]
- to:
- namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: kube-system}}
podSelector: {matchLabels: {k8s-app: kube-dns}}
ports: [{protocol: UDP, port: 53}, {protocol: TCP, port: 53}]
- to:
- namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: nats-system}}
ports: [{protocol: TCP, port: 4222}, {protocol: TCP, port: 8222}]
Each ServiceMonitor scrapes with the Prometheus pod’s token, which metrics.scraper.serviceAccount grants /metrics, and verifies the certificate against the Secret’s ca.crt as its Service’s DNS name, such as nats-operator-cluster-controller-metrics.nats-operator.svc.
The first two egress rules are the API server: the address of the kubernetes Service in default, and the addresses of its endpoints, which kubectl -n default get endpointslices -l kubernetes.io/service-name=kubernetes lists; both addresses above are examples. Port 4222 is the NATS cluster’s client port, which a NatsConnection below names, and 8222 its servers’ monitoring port, which the cluster controller reads.
Under the policy #
A NATS cluster and a stream, reconciled through those rules alone.
01-natscluster.yaml
apiVersion: cluster.nats.mikluko.io/v1beta1
kind: NatsCluster
metadata:
name: demo
namespace: nats-system
spec:
version: 2.15.0
replicas: 3
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
memory: 256Mi
jetstream:
volumeClaimTemplate:
spec:
storageClassName: standard
resources:
requests:
storage: 1Gi
kubectl apply -f https://nats-operator.io/docs/stories/12-metrics/01-natscluster.yaml01-status-natscluster-demo.yaml
status:
replicas: 3
readyReplicas: 3
conditions:
- type: Ready
status: "True"
reason: AllServersReady
02-natsconnection.yaml
apiVersion: nats.mikluko.io/v1beta1
kind: NatsConnection
metadata:
name: demo
namespace: nats-system
spec:
servers: ["nats://demo.nats-system.svc:4222"]
kubectl apply -f https://nats-operator.io/docs/stories/12-metrics/02-natsconnection.yaml02-natsstream.yaml
apiVersion: jetstream.nats.mikluko.io/v1beta1
kind: NatsStream
metadata:
name: events
namespace: nats-system
spec:
connectionRef:
name: demo
name: EVENTS
subjects: ["events.>"]
storage: File
replicas: 3
kubectl apply -f https://nats-operator.io/docs/stories/12-metrics/02-natsstream.yaml02-status-natsstream.yaml
status:
conditions:
- type: Ready
status: "True"
reason: Synced
- type: Synced
status: "True"
reason: MatchesSpec
Outside the policy #
The controllers reach nothing the policy omits. A NATS cluster in nats-outside, which no egress rule admits, comes up, as the cluster controller deploys it through the API server; a NatsConnection to it stays unready, as the JetStream controller’s dial is dropped.
01-natscluster-outside.yaml
apiVersion: cluster.nats.mikluko.io/v1beta1
kind: NatsCluster
metadata:
name: outside
namespace: nats-outside
spec:
version: 2.15.0
replicas: 1
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
memory: 128Mi
kubectl apply -f https://nats-operator.io/docs/stories/12-metrics/01-natscluster-outside.yaml01-status-natscluster-outside.yaml
status:
replicas: 1
readyReplicas: 1
conditions:
- type: Ready
status: "True"
reason: AllServersReady
02-natsconnection-outside.yaml
apiVersion: nats.mikluko.io/v1beta1
kind: NatsConnection
metadata:
name: outside
namespace: nats-system
spec:
servers: ["nats://outside.nats-outside.svc:4222"]
kubectl apply -f https://nats-operator.io/docs/stories/12-metrics/02-natsconnection-outside.yaml02-status-natsconnection-outside.yaml
status:
conditions:
- type: Ready
status: "False"
reason: ConnectFailed
Scraping by hand #
As the ServiceMonitor does, through a port-forward to the cluster controller’s metrics Service:
kubectl -n nats-operator get secret nats-operator-metrics-tls -o jsonpath='{.data.ca\.crt}' | base64 -d > ca.crt
kubectl -n nats-operator port-forward svc/nats-operator-cluster-controller-metrics 8080 &
curl --cacert ca.crt \
--resolve nats-operator-cluster-controller-metrics.nats-operator.svc:8080:127.0.0.1 \
-H "Authorization: Bearer $(kubectl -n monitoring create token prometheus)" \
https://nats-operator-cluster-controller-metrics.nats-operator.svc:8080/metrics