Fitness matrix
What's drop-in, supported, roadmap, and unsupported when running Kafka workloads on KubeMQ (T1–T4).
This page gives an honest, per-workload fitness verdict for a Kafka workload moving to KubeMQ — whether it's coming from Apache Kafka, Amazon MSK, or Confluent. For a verdict scoped to your own cluster, run the read-only assessor:
kmq assess kafka --bootstrap your-broker:9092 [--tls --sasl-mechanism scram-sha-256 --sasl-username … --sasl-password …]kmq assess kafka scans your cluster read-only — it never produces, commits, or
creates topics — and maps the same verdicts below onto your actual topics, configs, and
consumer groups. See Migrate from Kafka
for the full assess-then-migrate workflow.
Overview
Every row below carries a proof tier — the honesty guarantee that no claim is stronger than its evidence:
| Tier | Meaning |
|---|---|
| T1 | Proven with real clients and tools, on a cluster, in a shipped release |
| T2 | Proven on the current build; shipping in the pinned version |
| T3 | On the roadmap — designed, not yet proven |
| T4 | Not supported — with the real reason |
These tiers apply across Apache Kafka, Amazon MSK, and Confluent alike unless a row notes
otherwise. The verdicts below are for the Kafka connector in the kubemq-next image, the
image to install today. The
sections below group verdicts by outcome — 🟢 works, 🟡 caveat, 🟠 roadmap,
🔴 unsupported — and each row keeps its tier so you can see how proven a green checkmark
really is.
🟢 Works — drop-in fit
| Capability | Tier | Notes |
|---|---|---|
| Produce / consume, classic consumer groups, idempotent producers | T1 | The large majority of everyday Kafka usage — a straight repoint, no client changes. |
Compacted topics (cleanup.policy=compact) | T1 | Runs on the next storage engine, KubeMQ's default engine for Kafka. |
| Transactions / EOS (V1 wire surface) | T2 | Safe for at-least-once delivery plus basic exactly-once semantics — not a KIP-890 soundness guarantee. See the durability note below for the broader ack posture. |
| SASL/SCRAM + mTLS + ACLs + quotas | T1 | ACL authorization is coarse-grained — the Write permission covers both produce and destructive admin operations, a documented limitation rather than a roadmap gap. Quotas are a per-principal byte-rate baseline. |
| OAUTHBEARER / OIDC federated auth | T2 | Supported. The broker validates the client's OIDC bearer token and uses the token's sub claim as the authorization principal. Enforced only on the TLS/SASL_SSL listener — refused on plaintext. |
| Schema Registry — magic-byte serializer clients | T1 | Payloads are opaque to the broker; any Avro/Protobuf/JSON-Schema serializer works unchanged. |
Schema Registry — the real Confluent service (on _schemas) | T2 | Runs against KubeMQ; _schemas is a compacted topic on the next storage engine. Any node works — see below. |
| Kafka Connect — distributed workers | T2 | Config/offset/status internal topics plus a source→sink pipeline, proven across a broker failover. Any node works — see below. |
| Kafka Streams — stateful (aggregations) and stateless topologies | T2 | Compacted changelog restore across restart. Any node works — see below. |
Static membership (KIP-345, group.instance.id) | T1 | Full support. Static join bypasses MEMBER_ID_REQUIRED, a static rejoin skips rebalance, and a displaced/mismatched instance is fenced FENCED_INSTANCE_ID(82) across Join/Sync/Heartbeat/OffsetCommit/Leave. |
Any node works for admin requests. Topic creation, deletion, partition changes and config changes sent to a follower node are forwarded to the cluster leader, so the Schema Registry service, Kafka Connect and Kafka Streams can be pointed at any node.
Durability — the honest headline
With the next storage engine, a produce ack means the record has been fsynced to disk
on a quorum of nodes. Apache Kafka's default posture (acks=all, no per-write
flush.messages / flush.ms) acknowledges after quorum replication to page cache — fsync
timing is left to the OS — so a correlated power loss across replicas could lose
acknowledged writes that KubeMQ's default posture would not. Kafka can be configured to
fsync on every write, at a throughput cost — this is a comparison of defaults, not a
marketing claim.
The engine's ingest figure is approximately 183k messages/sec for a single replicated group. That's an engine-layer number, not a Kafka-wire throughput figure — benchmark the wire-level throughput and latency envelope on your own hardware.
🟡 Works with a caveat
| Capability | Tier | Caveat |
|---|---|---|
acks=0 (fire-and-forget) | T1 | On a cluster, an acks=0 produce sent to a non-leader node is dropped by design — there's no response channel to signal a redirect. Clients that follow cluster metadata and route to the leader are unaffected. |
| Share groups (KIP-932) — queue-style acquire/acknowledge consumption | T2 | Generally available. Certified with Java's KafkaShareConsumer (kafka-clients 4.3) and the Kafka 4.3 console tools on a single node and on three-node clusters, including a killed share coordinator with nothing lost and nothing accepted twice, plus soak runs of mixed Accept/Release/Reject traffic; franz-go works, and librdkafka's share consumer is an upstream preview. Caveats: a dead consumer's records come back only when their acquisition lock expires (30 seconds by default), the in-flight limit counts stored batches rather than records, and share.assignment.interval.ms is refused — full list. kmq migrate does not carry share-group start offsets; set them with kmq kafka share-groups reset-offsets before the first consumer joins. |
🟠 Roadmap, not proven
| Capability | Tier | Status |
|---|---|---|
Per-topic retention.bytes (size) enforcement | T3 | Size-based eviction is accepted but not enforced — on the roadmap; don't rely on it yet. Time-based retention.ms is enforced on the next storage engine: a leader-gated sweep runs every 10 seconds and evicts records older than the window for any finite positive value — retention.ms=-1, 0, or unset all pin the topic (never age-evicted), which differs from Apache Kafka, where retention.ms=0 evicts almost immediately. Proven on a 3-node SIGKILL crash-gate. See Durability & Retention. |
(History and committed-offset migration via kmq migrate is now shipped — see
History and committed-offset migration below.)
🔴 Not supported
| Capability | Tier | Why |
|---|---|---|
| Topics with > 256 partitions | T4 | KubeMQ caps a topic at 256 partitions. |
replication.factor above the node count | T4 | Refused at CreateTopics with INVALID_REPLICATION_FACTOR, as Kafka refuses a replication factor above its broker count. Any value up to the node count — the usual 3 on a three-node cluster, or -1 — is accepted, so topic scripts and the internal-topic defaults of Kafka Connect and Schema Registry work unchanged; every partition is replicated to every node regardless. |
max.message.bytes above the broker limit | T4 | Records larger than the broker-wide limit are rejected (MESSAGE_TOO_LARGE) — default 1 MiB, operator-configurable up to 1 GiB via Connectors.Kafka.MaxMessageBytes. The per-topic max.message.bytes config is echo-only and never raises the enforced limit. |
KIP-848 next-gen consumer groups (group.protocol=consumer) | T4 | The protocol keys aren't advertised; consumers fail client-side. Workaround: set group.protocol=classic — still the default for librdkafka and Java 3.x clients. |
| Kerberos / GSSAPI | T4 | No KDC integration — switch to SASL/SCRAM or mTLS. |
| Delegation tokens | T4 | Answered DELEGATION_TOKEN_AUTH_DISABLED(61) — switch those principals to SASL/SCRAM or mTLS. |
| Horizontal write scale (per-partition leadership) | T4 | Single replicated group — documented honestly, not a hidden limitation. |
Burrow / __consumer_offsets-tailing lag tools | T4 | Offsets live in an internal store, not a wire-visible compacted topic — use the OffsetFetch API and the connector's lag gauge instead. |
History and committed-offset migration
Migrating historical topic data and committed consumer-group offsets off an existing
Kafka cluster — so consumers resume where they left off, with no re-processing — is now
shipped via the kmq migrate
bridge, in four phases: assess → replicate → translate → cutover. Replication is
byte-identical, CreateTime-preserved, and partition-pinned, building an exact per-record
source→target offset map; cutover seeds each group's translated offsets via an empty-group
offset commit, and is fail-closed — a group committed past the durably (quorum-fsynced)
replicated watermark is refused rather than seeded past un-replicated data.
Proven on a 3-node next cluster: a SIGKILL mid-replication loses nothing, and a kill during
cutover never double-seeds. An oversized source record (over the 1 MiB target cap) blocks
and reports — it's never silently skipped.
Caveat to keep in view: cluster-level consumer-resume is proven with franz-go; Java/kcat/librdkafka resume is proven single-node, not yet across a cluster. See Migrate from Kafka for the full assess → replicate → translate → cutover workflow, per-source auth, and rollback.
Coming from Confluent
- OAuth / OIDC identity → KubeMQ speaks OAUTHBEARER against your existing IdP, over TLS.
- Schema Registry → run the real Confluent Schema Registry against KubeMQ (the
_schemastopic is a compacted topic on thenextstorage engine) — no re-platforming of your serializers. It can be pointed at any node (see above). - Kafka Streams → works today (stateful and stateless topologies; changelog restore across restart), pointed at any node.
- ksqlDB → not yet — a broker-side topic-creation fix for ksqlDB's control/query topics is in progress.
- Flink / Spark structured-streaming source/sink → on the roadmap, validated once the retention/eviction work lands.
See Also
Migrate from Kafka
Assess fit with kmq assess kafka, then move topics and consumer-group offsets with the kmq migrate tool.
Kafka
Overview of KubeMQ's Kafka connector — what works, how to assess fit, and how to migrate.
Configuration reference
The complete Connectors.Kafka.* settings, ports, and TLS/SASL options.
Was this page helpful?