Backup
A Backup is used for backing up topics from an EventHub, and offloading them to Storage. It is configured by creating a Backup resource in Kubernetes. The Kannika Armory Operator creates a StatefulSet that runs the backup process, based on the configuration in the Backup resource.
A Backup can have multiple Backup Streams which are used to configure which topics should be backed up.
Synopsis
Section titled “Synopsis”apiVersion: kannika.io/v1alphakind: Backupmetadata: name: my-backup labels: io.kannika/data-retention-policy: Delete # Optional: mark expired partitions for deletion io.kannika/data-retention-policy-delete-after: 30d # Optional: retention period before deletionspec: description: "Daily backup of production topics" # Optional: human-readable description source: "my-kafka-cluster" # EventHub resource to back up from sourceCredentialsFrom: # Optional: credentials for the source credentialsRef: # Reference to a Credentials resource name: my-source-credentials sink: "my-storage" # Storage resource to back up to sinkCredentialsFrom: # Optional: credentials for the sink credentialsRef: # Reference to a Credentials resource name: my-sink-credentials enabled: true # Set to false to pause the backup compression: algorithm: "zstd" # Optional: compression algorithm (none, zstd, gzip, bzip2, xz, deflate, brotli, lzma, zlib) quality: 3 # Optional: compression quality level segmentRolloverTriggers: size: "256MiB" # Optional: max segment file size before rolling over timeoutSeconds: 3600 # Optional: max time in seconds before rolling over workGroup: workers: 3 # Optional: number of workers for parallel processing seed: 42 # Optional: seed for consistent partition assignment lagMonitor: refreshInterval: "15s" # Optional: how often to query for latest offsets requestTimeout: "10s" # Optional: timeout for each ListOffsets request monitoring: enabled: true # Optional: enable monitoring for the backup resources: # Optional: resource requests/limits for the monitoring sidecar requests: cpu: "100m" memory: "128Mi" limits: cpu: "200m" memory: "256Mi" extraEnvVars: # Optional: extra environment variables for the backup pod - name: "RUST_LOG" value: "info" sourceAdditionalProps: # Optional: additional properties for the source connection consumer.max.poll.records: "1000" sinkAdditionalProps: # Optional: additional properties for the sink connection some.property: "value"
streams: # List of topics to back up - topic: "my-topic" enabled: true # Set to false to pause this stream
topicSelectors: # Optional: automatically import matching topics matchers: - name: literal: "exact-topic" - name: glob: "events.*" - name: regex: "^logs\\." - allOf: # Optional: combine conditions (allOf, anyOf, noneOf) - name: glob: "orders.*" - property: # Optional: match on a topic configuration property key: "cleanup.policy" operator: Contains value: "compact" excludeMatchers: # Optional: exclude topics from matching - name: literal: "do.not.backup" - property: key: "retention.ms" operator: Lt value: 3600000Backups can be managed using the kubectl command line tool,
and are available by the name backup or backups.
$ kubectl get backupsNAME STATUS AGEmy-backup Streaming 1sBackup Status
Section titled “Backup Status”A Backup can have the following statuses:
- Draft The Backup has no streams defined, and is not ready to be started.
- Paused The Backup is configured but it has not been started yet or it has been paused, and no data is being backed up.
- Initializing The Backup process is being created
- Streaming The Backup is running and backing up data to the storage.
- Error The Backup has failed, and no data is being backed up.
Additionally,
a Backup resource exposes a DeploymentReady Condition to report on the state of the underlying StatefulSet.
This condition’s value can either be ‘True’ or ‘False’,
indicating whether this Backup’s deployment is healthy.
The condition’s reason gives additional context and may be one of the following:
- DeploymentReady the deployment is healthy;
- DeploymentError the deployment encountered an error: check the status of this deployment to find out what happened;
- DeploymentDeleted somebody or something deleted the deployment and this should be a transient state;
- DeploymentStateUnknown the deployment state is unknown, perhaps due to an error talking with the kubernetes API.
Configuring a Backup
Section titled “Configuring a Backup”The following is an example of a Backup.
It configures a Backup which will back up two topics from the my-kafka-cluster EventHub to the my-bucket Storage.
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" streams: - topic: "magic.events" - topic: "pixie.dust"In this example:
-
A Backup named
backup-exampleis created, indicated by the.metadata.namefield. This name will become the basis for the StatefulSet which is created for this Backup. -
The Backup will connect to the
my-kafka-clusterEventHub to fetch data, indicated by the.spec.sourcefield. The Backup will write data to themy-bucketStorage defined in thespec.sinkfield. -
The
.spec.streamsfield contains a list of Backup Streams. A Backup Stream contains the configuration of each topic that will be backed up. In this case, two topics namedmagic.eventsandpixie.dustwill be backed up.
Automatically importing topics
Section titled “Automatically importing topics”It is possible to configure a backup so that it automatically adds topics present on a cluster if they match a condition:
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" streams: - topic: "magic.events" topicSelectors: matchers: - name: literal: "spells.proper" - name: glob: "enchantments.*" - name: regex: "^curses\\." - allOf: - name: glob: "potions.*" - property: key: "cleanup.policy" operator: Contains value: "compact"In this example, we have changed the streams in our Backup definition:
The “magic.events” topic is still present,
but we added topic “matchers” under the spec.topicSelectors property
that will watch the cluster for new (or existing) topics matching any one of their rules:
-
the topic called
spells.properif it is present, or should it appear at some point; -
any new (or existing) topic matching the glob pattern
enchantments.*; -
any new (or existing) topic matching the regular expression
^curses\\.; -
any new (or existing) topic matching the glob pattern
potions.*that is also compacted, see Matching on topic properties and Combining conditions.
Any dynamically added topic via the topicSelectors property
will be backed up using the same configuration options (compression, rollover size, etc)
as the ones explicitly defined in spec.streams.
Excluding topics
Section titled “Excluding topics”It is possible to exclude topics from being backed up using excludeMatchers.
The excludeMatchers rules take precedence over the matchers rules.
Topics defined in spec.streams are not affected by excludeMatchers rules,
and will be backed up regardless of any excludeMatchers rule.
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" streams: - topic: "temp.included" # This topic will be backed up topicSelectors: matchers: - name: glob: "*" # Match all topics excludeMatchers: - name: literal: "do.not.backup" # Exclude this specific topic - name: regex: "^internal\\..*" # Exclude topics that start with "internal." - name: glob: "temp.*" # Exclude topics that start with "temp." - allOf: # Exclude "archive." topics that keep data for less than a day - name: glob: "archive.*" - property: key: "retention.ms" operator: Lt value: 86400000In this example,
we are matching all topics on the cluster using the matchers rule with a glob of "*".
However, we are excluding the following topics from being backed up:
- The topic named
do.not.backup - Any topic matching the regex
^internal\\..* - Any topic matching the glob pattern
temp.*, except for thetemp.includedtopic which is explicitly defined inspec.streamsand will be backed up regardless of theexcludeMatchersrules. - Any topic matching the glob pattern
archive.*that keeps its data for less than a day.
Matching on topic properties
Section titled “Matching on topic properties”Besides its name,
a topic can be matched on the properties of its configuration,
such as cleanup.policy or retention.ms.
A property condition takes the key of the property,
an operator,
and depending on the operator a value or a list of values.
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" topicSelectors: matchers: - property: key: "cleanup.policy" operator: Contains value: "compact" excludeMatchers: - property: key: "retention.ms" operator: Lt value: 3600000In this example,
every topic whose cleanup.policy contains compact is backed up,
unless its retention.ms is lower than one hour.
The following operators are available:
| Operator | Operands | Matches when the property |
|---|---|---|
Equals | value | is set and equal to the value |
NotEquals | value | is not set, or not equal to the value |
Contains | value | is set and contains the value as a substring |
NotContains | value | is not set, or does not contain the value as a substring |
In | values | is set and equal to one of the values |
NotIn | values | is not set, or not equal to any of the values |
Exists | none | is set |
DoesNotExist | none | is not set |
Gt | value | is set to an integer greater than the value |
Lt | value | is set to an integer lower than the value |
In and NotIn also accept a single value instead of a list of values.
Gt and Lt require an integer value,
and never match a property that is not set or is not an integer.
Values are compared as text,
so a boolean property is matched with a quoted value such as "true".
To match on properties,
the Backup describes the configuration of the topics on the cluster.
This requires the DescribeConfigs permission on those topics,
in addition to the permissions that are needed to back them up.
Topic configurations are only described when a selector uses a property condition,
so the DescribeConfigs permission is only needed then.
Without it, topics are not picked up, and the Backup reports “Could not read the topic configuration from the event hub”.
Topics that are already being backed up keep running.
When no topic is being backed up, the Backup stops once the retry budget from Intervals and timeouts runs out.
The configuration of topics that did not match is described again every 10 minutes, so a topic whose configuration starts matching later on is picked up without restarting the Backup. A topic that is already being backed up keeps being backed up, even when its configuration no longer matches. See Intervals and timeouts to change the interval.
Combining conditions
Section titled “Combining conditions”Conditions can be combined with the allOf, anyOf and noneOf groups:
allOfmatches when all of its conditions match;anyOfmatches when at least one of its conditions matches;noneOfmatches when none of its conditions match.
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" topicSelectors: matchers: - allOf: - anyOf: - name: glob: "letters.*" - name: glob: "photos.*" - property: key: "retention.ms" operator: Equals value: "-1" - noneOf: - name: glob: "*.draft"In this example,
every letter and photo sealed in the time capsule is backed up:
topics whose name matches letters.* or photos.*,
and that keep their records forever (retention.ms is -1).
Drafts are left out.
Groups can be nested up to 3 levels deep,
and must contain at least one condition.
Each entry in matchers and excludeMatchers,
and each condition inside a group,
holds exactly one of name, property, allOf, anyOf or noneOf.
Intervals and timeouts
Section titled “Intervals and timeouts”A Backup with topicSelectors lists the topics on the cluster at a regular interval to find topics that match.
The timing can be changed with the following environment variables in spec.extraEnvVars:
| Environment variable | Default | Effect |
|---|---|---|
KANNIKA_DISCOVERY_REFRESH_INTERVAL_SECS | 120 | The time between two listings of the topics on the cluster. |
KANNIKA_DISCOVERY_LIST_TIMEOUT_SECS | 30 | How long a single listing may take. |
KANNIKA_DISCOVERY_RETRY_BUDGET_SECS | 300 | How long discovery may keep failing before the Backup gives up, when no topic is being backed up. |
KANNIKA_DISCOVERY_DESCRIBE_CONFIG_INTERVAL_SECS | 600 | The time between two describes of the topic configurations, when matching on properties. |
KANNIKA_DISCOVERY_DESCRIBE_CONFIG_TIMEOUT_SECS | 30 | How long a single describe of topic configurations may take. |
Each value is a whole number of seconds. A value that cannot be parsed falls back to the default.
The following example sets all of them:
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: backup-examplespec: source: "my-kafka-cluster" sink: "my-storage" extraEnvVars: - name: "KANNIKA_DISCOVERY_REFRESH_INTERVAL_SECS" # List the topics every minute value: "60" - name: "KANNIKA_DISCOVERY_LIST_TIMEOUT_SECS" # Allow a listing to take up to 1 minute value: "60" - name: "KANNIKA_DISCOVERY_RETRY_BUDGET_SECS" # Keep retrying a failing discovery for 10 minutes while no topic is being backed up value: "600" - name: "KANNIKA_DISCOVERY_DESCRIBE_CONFIG_INTERVAL_SECS" # Describe the topic configurations every 5 minutes value: "300" - name: "KANNIKA_DISCOVERY_DESCRIBE_CONFIG_TIMEOUT_SECS" # Allow a describe to take up to 1 minute value: "60" topicSelectors: matchers: - property: key: "cleanup.policy" operator: Contains value: "compact"Pausing and resuming
Section titled “Pausing and resuming”It is possible to pause a Backup, and resume it later.
When a Backup is paused,
the associated StatefulSet will be scaled down to 0.
Pausing and resuming a Backup
Section titled “Pausing and resuming a Backup”You can pause a Backup by setting the .spec.enabled field to false.
The following is an example of a Backup which is disabled (paused).
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: paused-backup-examplespec: source: "my-kafka-cluster" sink: "my-bucket" enabled: false streams: - topic: "paused.events"A Backup can be resumed again by setting the spec.enabled field to true again (or by removing the field).
Pausing and resuming a Backup Stream (Topic)
Section titled “Pausing and resuming a Backup Stream (Topic)”A Backup Stream can be paused by setting the enabled field to false of a Backup Stream.
The following is an example of a Backup Stream which is paused.
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: paused-backup-stream-examplespec: source: "my-kafka-cluster" sink: "my-bucket" enabled: false streams: - topic: "disabled.events" enabled: false # Disables this streamIf a topic satisfies one of the topicSelectors matchers,
then it is possible to pause the Backup for this particular topic by adding it to the spec.streams list
with the enabled flag set to false,
as shown in the following example:
apiVersion: kannika.io/v1alphakind: Backupmetadata: name: paused-backup-stream-examplespec: source: "my-kafka-cluster" sink: "my-bucket" enabled: true streams: - topic: "__consumer_offsets" enabled: false topicSelectors: matchers: - name: glob: "*"Here,
we the topicSelectors rules match all possible topic names,
but we’ve excluded the __consumer_offsets topic by setting enabled.false.

