Chaos engineering — preview and safe setup
Chaos engineering deliberately disrupts a target system so you can test its
recovery. Start in a staging cluster, choose a narrow target, and get approval
from the people responsible for that cluster. This guide explains the available
faults, the required opt-in settings, and the safeguards to check before a run.
What preview means for an operator
Fault injection can stop pods, change network behaviour, or consume resources.
Some Kubernetes faults require elevated privileges in the target cluster. Before
using them, review the account and cluster permissions, the target selection,
and your recovery procedure. The current operator does not enforce image
signatures or validate experiment resources through an admission webhook.
Apply your cluster’s own policies and check that they reject an unsafe test
resource before relying on them. The CLI requires an explicit opt-in for
fault-injecting commands.
Available faults
The preview supports the following fault types:
pod_kill/pod_kill_random— terminate one or more target pods.network_partition— block all ingress/egress for a pod.network_delay— add latency to a pod’s network traffic.network_loss— drop a percentage of packets.deployment_restart— rolling restart on a deployment.scale— scale a deployment to N replicas, then restore.node_drain— cordon + drain a node.resource_stress— CPU / memory / IO pressure.dns_disruption— break DNS resolution for a target domain.io_chaos— corrupt or delay file-system operations.time_chaos— skew the pod’s clock.
These are surfaced through:
- The Web UI Chaos dashboard.
mockarty-cli chaos run/mockarty-cli chaos preset run.- The TUI chaos screen (
mockarty-cli tui chaos). - The SDK
Chaosclients (Go / Python / Java). - Test Plans (when chaos items are added — gated by the licence’s chaos
feature).
What’s gated
The opt-in gate sits on the CLI side and applies to every
fault-injecting subcommand:
mockarty-cli chaos run …mockarty-cli chaos abort …mockarty-cli chaos preset run …mockarty-cli chaos cluster add …(registers a new K8s connection)mockarty-cli chaos operator install …/uninstall …
These commands fail closed without an opt-in. The gate is independent
of the licence: it’s an extra defensive check that runs before the
licence is consulted.
The following read-only commands work without opt-in so you can audit
prior runs and inspect the cluster state:
mockarty-cli chaos listmockarty-cli chaos get <id>mockarty-cli chaos report <id>mockarty-cli chaos preset listmockarty-cli chaos operator statusmockarty-cli chaos cluster list/cluster info
How to opt in
There are two equivalent ways. Use whichever fits your workflow.
1. CLI flag (per invocation)
mockarty-cli chaos run \
--type pod_kill \
--namespace app \
--selector app=web \
--duration 5m \
--i-know-this-is-preview
The flag is persistent on the chaos parent command — every subcommand
inherits it. There is no abbreviated form; the full phrase is intentional
so it lands in shell history and code review as a deliberate decision.
2. Environment variable (per session / CI)
export MOCKARTY_CHAOS_PREVIEW=1
mockarty-cli chaos run --type pod_kill --namespace app --selector app=web --duration 5m
Accepted true-values: 1, true, yes, on (case-insensitive). Any
other value (including the empty string) is treated as “not acknowledged”.
In CI, set MOCKARTY_CHAOS_PREVIEW=1 only for the job that runs chaos
experiments. Other jobs can then inspect results without being able to
start a fault by accident.
What the banner looks like
If you forget to opt in, you’ll see:
Chaos engineering is currently in pre-GA preview.
Fault-injection subcommands are gated behind an explicit opt-in:
--i-know-this-is-preview (CLI flag)
MOCKARTY_CHAOS_PREVIEW=1 (environment variable)
Read-only operations (list, get, report, preset list, operator
status) work without opt-in so you can audit prior experiments.
Full details: https://mockarty.ru/docs/chaos-pre-ga
Exit code is 3, distinguishable from licence errors (which exit 2).
Server-side licence is a separate check
The CLI opt-in and the server’s Chaos & Reliability licence check are
separate. Without the required licence and assigned seat, the admin node
returns 403 regardless of the CLI flag:
| CLI opt-in | Chaos entitlement | Result |
|---|---|---|
| no | no | CLI exits with the preview banner |
| no | yes | CLI exits with the preview banner |
| yes | no | Server returns 403 feature not licensed |
| yes | yes | Command runs |
Read-only commands skip the CLI gate but still respect the server-side
licence (with the standard historical-data fallback — if your workspace
had chaos last year you can still read prior experiments after lapse).
Maturity
Chaos engineering is in preview. The listed fault types are available,
and the CLI opt-in applies to fault-injecting commands. Check the current CLI
help and your installation’s permissions before changing a production script.
Licence and audit checks
For regulated environments (financial services, gov-cloud, healthcare):
- Experiment history shows runs and their results. If you need an audit trail
or SIEM evidence, verify the exact events recorded by your installation
before relying on it. The operator’s event names alone do not establish
that those events reach the audit log. - Chaos & Reliability is a seat-pool licence module; users without a
seat can’t initiate experiments even with the CLI gate opt-in.
FAQ
Why a separate flag instead of the licence?
The licence controls who may use chaos engineering. The preview flag is a
separate acknowledgement before the CLI sends a fault-injecting command.
Can I disable the gate globally on a dev machine?
Yes — set MOCKARTY_CHAOS_PREVIEW=1 in your shell’s RC file. The CLI
re-reads the env on every invocation.
What should a CI job set?
Pass --i-know-this-is-preview or set MOCKARTY_CHAOS_PREVIEW=1 for the
current CLI. Keep the target and licence checks in the job as well.
Where can I see the result of a run?
Use the experiment history and its report. If your process requires an audit
record, verify it separately in your installation’s audit log and SIEM.
Before using a production cluster
Limit the namespaces and resources the operator may change, keep a tested
rollback procedure, and use your organisation’s image policy. The operator’s
Helm values for admission and image signatures do not activate those checks
in the current runtime. Do not treat the presence of these settings or a
successful Helm install as proof that the cluster enforces them.