KubeMQ
ConfigureReference

Deployment & High Availability

Kubernetes operator packaging — image, storage, resources, health, scheduling, Service exposure, and clustered high availability.

These settings cover how KubeMQ is packaged and run on Kubernetes — the container image, persistent storage, resource requests/limits, health probes, node scheduling, and Service exposure — plus the high-availability replica controls.

Unlike the rest of this reference, these fields do not exist in the server's config.yaml: they are typed on the KubemqCluster CRD and consumed by the operator, which renders the StatefulSet, Services, and PersistentVolumeClaim. The pattern is therefore inverted — the Helm/CRD path is the real one, and the Docker column is — (Kubernetes-only) except where the operator translates a CRD field into a pod env var (shown as · env VAR). On Docker single-node, the equivalent concerns are docker run flags (-p, -v, --cpus / --memory) covered in the Docker guide.

Kubernetes packaging

Container image

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
ImagestringThe server image of the installed operator releaseimage reference— (K8s-only)spec.image.imageLeave it unset: the cluster then runs the server image its operator release sets (Artifacts and versions; each release's image is in the release notes). Resolution order: spec.image.image → operator env RELATED_IMAGE_KUBEMQ_CLUSTER → built-in fallback (config/image.go:8,18-30). No server env binding.
Pull policystring (enum)AlwaysIfNotPresent / Always / Never—spec.image.pullPolicyCRD pattern (IfNotPresent|Always|Never); empty is coerced to Always (config/image.go:14,34-37).
Image pull secretsstring[][]Secret names—— (chart value imagePullSecrets[])The chart value imagePullSecrets applies to the operator pod only; it never reaches the cluster resource. For server pods pulling from a private mirror, see Install air-gapped.

Storage (volume)

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Volume sizestringunset → ephemeral (no PVC)k8s quantity (e.g. 50Gi)—spec.volume.sizeWhen set, the operator renders a volumeClaimTemplates PVC (ReadWriteOnce) mounted at ./kubemq/store. When unset/empty, the pod has no persistent volume and the store is ephemeral (deployment.go:78-81, deployment/statefulset.go:108-124). There is no built-in 10Gi default in code. Required for durable data on the next storage engine — such a cluster with no spec.volume.size set raises the EphemeralNextStore warning (its durable data would otherwise live on ephemeral container storage).
Storage classstring"" → cluster defaultStorageClass name—spec.volume.storageClassOnly consulted when size is set; empty renders a blank storageClassName: → the cluster's default StorageClass (config/volume.go:8).

spec.store.path on the next storage engine: the operator rejects a leading /. The operator itself supplies the mount-rooted absolute path for the PVC, so a spec.store.path value starting with / is rejected outright — the operator, not the server, owns that absolute path. This is the inverse of the server-layer behavior documented on Storage & Queues: the next storage engine's server process honors an absolute StorePath verbatim (it's the legacy engine that rewrites a leading / to ./). Keep the two facts distinct — "the operator rejects a leading / in the CR field" and "the server on the next storage engine process honors an absolute StorePath" are both true, at different layers.

Resources

Each field is rendered into the pod's resources: block only when non-empty; there are no defaults (config/resources.go:32-47).

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
CPU limitstringunsetk8s CPU quantity (e.g. 2, 500m)—spec.resources.limitsCpuPod resources.limits.cpu.
Memory limitstringunsetk8s memory quantity (e.g. 2Gi)—spec.resources.limitsMemoryPod resources.limits.memory.
Ephemeral-storage limitstringunsetk8s quantity—spec.resources.limitsEphemeralStoragePod resources.limits.ephemeral-storage.
CPU requeststringunsetk8s CPU quantity—spec.resources.requestsCpuPod resources.requests.cpu.
Memory requeststringunsetk8s memory quantity—spec.resources.requestsMemoryPod resources.requests.memory.
Ephemeral-storage requeststringunsetk8s quantity—spec.resources.requestsEphemeralStoragePod resources.requests.ephemeral-storage.

Health probe

The liveness probe is off by default and only injected when enabled: true (config/health.go:46-60).

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Enabledboolfalsetrue / false· env API_BIND_ADDRESS=0.0.0.0 (side-effect when true)spec.health.enabledWhen true, the operator adds a livenessProbe httpGet /health on the API port and sets API_BIND_ADDRESS=0.0.0.0 so the kubelet (pod IP) can reach the probe — the API binds 127.0.0.1 otherwise (config/health.go:48-60).
Initial delay (s)int325any; ≤0 → 5—spec.health.initialDelaySecondsconfig/health.go:38-40.
Period (s)int3210any; ≤0 → 10—spec.health.periodSecondsconfig/health.go:35-37.
Timeout (s)int325any; ≤0 → 5—spec.health.timeoutSecondsconfig/health.go:32-34.
Success thresholdint321any; ≤0 → 1—spec.health.successThresholdconfig/health.go:29-31.
Failure thresholdint3212any; ≤0 → 12—spec.health.failureThresholdconfig/health.go:41-43.

/health (liveness) vs /ready (readiness). The table above covers the optional /health liveness probe (spec.health.enabled, off by default). Readiness is separate: for a clustered next-engine deployment, the operator wires a readinessProbe against a /ready endpoint on the API port — it holds the pod not-Ready until cluster quorum forms, so traffic only reaches pods that can actually serve. The response code varies by mode, so it isn't listed here.

Node scheduling

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Node selectorsmap[string]string{}node label key/value pairs—spec.nodeSelectors.keysRendered as the pod nodeSelector; an empty map applies no constraint (config/node_selectors.go:9-24).

Service exposure & interface toggles

The three built-in interface Services — gRPC (50000), REST/WebSocket (9090), and API/dashboard (8080) — each expose a Service type, optional NodePort, custom port, and a disable toggle. These HTTP/interface Services are opt-out via disabled: true.

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
gRPC disabledboolfalsetrue / false· env CONNECTORS_GRPC_ENABLE=falsespec.grpc.disabledDeletes the -grpc Service and disables the listener (config/grpc.go:64-66, deployment.go:203-206).
gRPC Service typestring (enum)ClusterIPClusterIP / NodePort / LoadBalancer—spec.grpc.exposeEmpty → ClusterIP (config/grpc.go:56); CRD pattern-validated.
gRPC NodePortint320 (auto-assigned)30000–32767—spec.grpc.nodePortApplied only when expose: NodePort and value > 0 (config/grpc.go:77-81).
gRPC portint3250000port· CONNECTORS_GRPC_PORTspec.grpc.portMoves the container/Service/target port and emits the env (config/grpc.go:71-75).
REST disabledboolfalsetrue / false· env CONNECTORS_REST_ENABLE=falsespec.rest.disabledThe -rest Service (9090) is shared by REST/MCP/Agents/CE — the operator removes it only when all four are disabled (deployment.go:99-120).
REST Service typestring (enum)ClusterIPClusterIP / NodePort / LoadBalancer—spec.rest.exposeEmpty → ClusterIP (config/rest.go:71).
REST NodePortint320 (auto)30000–32767—spec.rest.nodePortApplied only when expose: NodePort and > 0 (config/rest.go:91-95).
REST portint329090port· CONNECTORS_REST_PORTspec.rest.portconfig/rest.go:85-88.
API disabledboolfalsetrue / false· env API_ENABLE=falsespec.api.disabledDeletes the -api Service (config/api.go:65-66, deployment.go:89-92).
API Service typestring (enum)ClusterIPClusterIP / NodePort / LoadBalancer—spec.api.exposeEmpty → ClusterIP (config/api.go:58).
API NodePortint320 (auto)30000–32767—spec.api.nodePortApplied only when expose: NodePort and > 0 (config/api.go:79-83).
API portint328080port· API_PORTspec.api.portconfig/api.go:73-76.

spec.statefulsetConfigData is a one-way door for every other packaging field. It does not merge with, patch, or extend the StatefulSet the operator builds — it replaces the body outright. The moment it is set, spec.image, spec.volume, spec.resources, spec.health, spec.nodeSelectors, spec.podAntiAffinity and spec.terminationGracePeriodSeconds stop having any effect on the pod, silently: they stay in the CR, they still validate, and they change nothing. You now own the whole pod spec, including the parts the operator was maintaining for you across upgrades. Use it only when a field you need genuinely has no typed equivalent, and expect to re-check it on every operator upgrade.

Only the exposure subset of spec.grpc / spec.rest / spec.api lives here. Their tuning fields (buffer/body limits, gRPC reflection, REST read/write timeouts, CORS, API allow-origins, and the API-auth block) are documented on the Interfaces and Security pages.

License

Set exactly one of the four fields: the chart and the CRD refuse none or more than one. Which license you need is on Plans compared. On the Helm path, Install on Kubernetes creates the Secret messaging-license that spec.licenseKeySecretRef reads; kmq names its Secret messaging-license- plus eight hexadecimal characters.

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
License keystringunset43-character license key— · env KUBEMQ_LICENSE_KEYspec.licenseKeyLiteral key, stored in the custom resource and in Helm values; prefer spec.licenseKeySecretRef. Rendered into the pod environment. The CRD rejects an object without exactly one license* field (exactly one license source is required). See Install on Kubernetes.
License key from Secret{name, key}unsetexisting Secret name + data key— · env KUBEMQ_LICENSE_KEYspec.licenseKeySecretRefRecommended. The operator reads the Secret's data key (default licenseKey) at each reconcile and copies the value into the pod-environment Secret it owns; your Secret is only ever read, and the value never appears on the CR or in Helm values.
License filestringunsetarmored license file contents— · env KUBEMQ_LICENSE_DATAspec.licenseFileOffline license file as a literal. The operator verifies it (signature, key id, not revoked, offline kind) before creating or updating the StatefulSet; a file that fails sets condition LicenseInvalid and the operator keeps checking until a valid file is supplied. See Install air-gapped.
License file from Secret{name, key}unsetexisting Secret name + data key— · env KUBEMQ_LICENSE_DATAspec.licenseFileSecretRefOffline license file from a Secret (data key defaults to licenseFile); read and copied like the key, and verified like spec.licenseFile. The operator watches the referenced Secret, so a corrected file is reconciled at once.

Advanced / other top-level spec fields

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
StatefulSet overridestring (YAML)unsetStatefulSet YAML fragment—spec.statefulsetConfigData⚠️ Escape hatch — when set, replaces the operator-generated StatefulSet body with your fragment (deployment/statefulset.go:126-135,268-271). All other packaging fields above are ignored for the STS. See the warning below.
Config data (OIDC)stringunsetOIDC block—spec.configDataOIDC-only passthrough — not a generic config.yaml. See Security.
Env overlaymap[string]string{}env key/value pairs—spec.envLast-wins overlay onto the pod ConfigMap — the general escape hatch for any server env key the CRD doesn't type. Operator-computed identity keys are rejected: STORE_ENGINE, CLUSTER_ENABLE, CLUSTER_NAME, CLUSTER_ROUTES, API_BIND_ADDRESS, CHECKSUM, POD_NAME, and any CLUSTER_REPLICATION_* key — not a CLUSTER_* wildcard (e.g. CLUSTER_PORT is still allowed) — plus the licensing variables (KUBEMQ_LICENSE_KEY, KUBEMQ_LICENSE_FILE, KUBEMQ_LICENSE_DATA, KUBEMQ_LICENSE_CACHED_LEASE) and KUBEMQ_API_OPERATOR_TOKEN, which reach the pods only from operator-owned Secrets (rejected with a ReconcileError condition, nothing silently stripped). Not for secrets.
Env from Secretsstring[][]existing Secret names—spec.envFromSecretsProjects existing Secret(s) as pod env (envFrom) — the general Secret-envFrom escape hatch; Secret values never transit the operator. This is the mechanism behind Kafka SASL credentials.

High availability

High availability on Kubernetes is multiple replicas managed by the operator — not the Docker cluster.* block.

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Replica count*int3233 or more—spec.replicasSupported Kubernetes clusters require at least three replicas. Nil or 0 defaults to 3; a request below 3 is refused before workload creation. Backs the scale subresource (specpath=.spec.replicas). The operator wires CLUSTER_* and the 5228 cluster port. Once the engine is established as next, spec.replicas is immutable — a CEL rule plus an operator fail-closed guard reject any change (see the engine-establishment guard below).

spec.replicas becomes immutable once the engine is next — and next is what a new cluster gets. This is a one-way door, armed at creation on a default install, not an edge case for people who opted into next. Once the engine is established as next, a CEL rule on the CRD and an operator fail-closed guard both reject any change to the replica count — up or down. helm upgrade --set replicas=5 fails; so does editing the CR.

Choose the replica count before you create the cluster. 3 is the minimum supported count, and standalone: true is rejected. See Requirements and supported setups for what the cluster network must provide. Three is also the smallest count that gets a disruption budget; see Disruption budget & pod spread. Changing it afterwards means creating a new cluster and migrating, so decide with the ceiling in mind rather than the starting load.

Disruption budget & pod spread

Both default to on for a multi-replica cluster — the operator creates them without being asked, and these fields exist to tune or disable them.

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Disruption budget enabledbooltruetrue / false—spec.podDisruptionBudget.enabledSet false to create no PodDisruptionBudget.
Minimum availableint32raft majority (replicas/2 + 1)≥ 1—spec.podDisruptionBudget.minAvailableDefaults to 2 of 3, 3 of 5. Overriding it downward trades quorum safety for drainability.
Anti-affinity enabledbooltruetrue / false—spec.podAntiAffinity.enabledSpreads replicas across nodes.
Anti-affinity requiredboolfalsetrue / false—spec.podAntiAffinity.requiredfalse (default) = preferred spread; true = required. See the trade below.

Use an odd replica count, 3 or more. A request below three replicas is refused before workload creation. A budget is created from 3 replicas up.

An even replica count gets an EvenReplicaCount Warning event: the majority of 4 is 3, the same single-failure tolerance as 3, at the cost of an extra replica and another copy of the data.

These two events are your only signal — nothing else reports that no budget was created:

kubectl get events -n <namespace> --field-selector reason=UnsupportedReplicaCount

The anti-affinity trade is real — choose it deliberately. required: false (the default) means pods always schedule, but two replicas may share a node, so a single node loss can still cost quorum. required: true guarantees the spread, but a replica stays Pending while there are fewer schedulable nodes than replicas.

Shutdown grace period

SettingTypeDefaultValid valuesDocker (config.yaml key · env var)Helm/CRD pathNotes
Termination grace period (s)int32unset → Kubernetes' 30≥ 1— (docker stop -t)spec.terminationGracePeriodSeconds · env KUBEMQ_TERMINATION_GRACE_PERIOD_SECONDSThe pod's shutdown budget. The operator sets the pod spec and passes the same number to the server — they always move together.

This setting governs message durability on shutdown, not just tidiness. The server splits the grace period between connector teardown — which requeues in-flight messages that would otherwise be lost — and store shutdown.

Kubernetes does not expose terminationGracePeriodSeconds to the container, so the operator both sets the pod spec and passes the same value as KUBEMQ_TERMINATION_GRACE_PERIOD_SECONDS. Left unset, Kubernetes uses 30s and the server assumes 30s.

The concrete numbers: at the default 30-second grace the requeue backstop is 1 second. 45 seconds funds the full 12-second backstop. Set at least 45 on any cluster where losing in-flight messages on a pod roll matters, and higher still if your connectors hold long in-flight batches.

cluster-values.yaml
terminationGracePeriodSeconds: 45

Connector Services & ports (operator-exposed). The pod always publishes container ports for every wire connector — MQTT 1883/8883/8083, AMQP 5672/5671, STOMP 61613/61614, AWS 4566, GCP 8085, Kafka 9092/9093 (deployment/statefulset.go:70-102, deployment/service.go:192-256). The matching ClusterIP Service is created only when that connector is enabled — otherwise the operator prunes it (deployment.go:122-199). Kafka and AMQP 0.9.1 are on by default, so their Services (9092/9093, 5672/5671) are created by default and removed only with spec.kafka.enabled: false / spec.amqp.enabled: false. The other five wire connectors (MQTT, AMQP 1.0, STOMP, AWS, GCP) are opt-in (enabled: true, default off) and get a Service only once enabled; the HTTP interfaces above (gRPC/REST/API) remain opt-out (disabled: true).

status.engine. The KubemqCluster status subresource exposes status.engine (legacy / next / auto) — the engine established for this cluster by the guard below. Read it alongside status.replicas when auditing the replicas-freeze caveat above: once status.engine reports next, spec.replicas can no longer change.

auto is not an engine. It means the cluster delegated the choice to the server, which resolves it at boot from the store directory. On such a cluster status.engine and the established-engine annotation both read auto, and the engine actually running is reported only in the server's boot NOTICE in the pod log — see Which engine am I actually on?.

Engine-establishment guard

The operator records the live persistence engine in the next.kubemq.io/established-engine annotation (legacy / next / auto) — once present, this annotation is authoritative on an established cluster, outranking spec.store.engine for every guard decision. A mismatch between the two is refused, not silently applied. Derivation runs in order: (1) the annotation, if present; (2) else the explicit spec.store.engine; (3) else the pod ConfigMap's STORE_ENGINE key; (4) else, if a StatefulSet or retained PVC already exists, the engine is unknown and the operator refuses to guess — set spec.store.engine explicitly or delete the retained PVCs; (5) else the cluster has no engine on record and the operator delegates: it writes STORE_ENGINE=auto and the server resolves the engine at boot from the store directory (clean ⇒ next).

A fresh cluster gets auto, not a pinned engine — the operator does not override the server's clean-store default. An established cluster keeps its named engine and nothing rolls. Setting spec.store.engine later on an auto cluster is allowed, because auto records a delegation rather than an established engine — pin the engine the server actually resolved; a pin that disagrees with the data on disk is refused by the server at boot, and the server never deletes a datadir.

The operator never changes a live cluster's engine. Two advisory CEL rules on the CRD back this: spec.store.engine is immutable once set (born-one-mode), and spec.replicas is immutable once the engine is established as next.

This operator-side derivation runs before the server-side probe documented on Storage Engines and the Kafka callout on Connectors — two layers of one decision: the operator decides what the pod's env will say before the pod ever boots, and when that env says auto, the server-side probe is what actually decides once it does.

Example

Publish the gRPC port on each target. On Kubernetes you set the Service exposure and node port; on Docker you publish the port with -p. This is a single-setting snippet — see the Kubernetes guide for complete, runnable configurations.

docker run -d \  --pull always \  --platform linux/amd64 \  --name kubemq \  --hostname kubemq \  -p 127.0.0.1:50000:50000 \  -e STORE_ENGINE=next \  -e STORE_NEXT_ACK_POLICY=strict \  -e STORE_STORE_PATH=/kubemq/store \  -e API_BIND_ADDRESS=0.0.0.0 \  -v kubemq-data:/kubemq/store \  europe-docker.pkg.dev/kubemq/images/kubemq-next:latest
cluster-values.yaml
grpc:
  expose: NodePort
  nodePort: 32000

For the full install flow (custom resource definitions → operator → cluster) and cluster-values.yaml mapped to the KubemqCluster specification, see Install on Kubernetes.

Was this page helpful?

On this page