Advanced Test Plans
Companion guide to Test Plans. This page covers the
execution-layer features:
- typed dependency gates (
on_success/on_fail/always/on_skipped) - per-item retry, timeout,
criticalandcontinueOnErrorflags - stable
itemUiddependency keys - the live run view (Graph / Timeline / Items / Logs / Artifacts / Report)
- six report formats and run-to-run comparison
- the optional distributed runner path for plan items
- the entity-search backend that powers every in-UI picker
About URLs in examples: we use
http://localhost:5770as the default
Mockarty address. Replace it with your own installation URL if different.
See Tips & Useful Features.
Related pages: Test Plans ·
Test Plans API Cookbook ·
Test Plans in CI/CD ·
Entity Search ·
SDK Guide
Execution modes
Every plan is saved with one of three modes (executionMode field).
| Mode | When to use | How it runs |
|---|---|---|
fifo |
Legacy sequential plans; one step at a time. | Items run in ascending order; a failure short-circuits remaining items (unless overridden). |
parallel |
Independent items that can all start at once. | All items dispatched concurrently; no dependency edges honoured. |
dag |
Typed dependencies with gates (default for new plans). | Items run in dependency order; each gate decides when its item becomes eligible. |
If you omit executionMode, the server auto-detects dag when any item
carries typed gates, otherwise fifo. For backward compatibility the
older schedule value (parallel / dag) is still accepted.
Items and dependencies
ItemUID — stable reference key
Every item in a plan carries an itemUid (UUID). It is generated server-side
when the plan is first saved and stays stable across edits. Dependency edges
point at itemUid, not at the external entity UUID (refId), so:
- you can add the same collection / config twice (e.g. smoke + regression +
cleanup) without ambiguity; - swapping the underlying entity (via
refIdedit) does not break the plan
graph; - renaming or deleting an entity only affects that one item, not siblings.
Legacy plans that pre-date itemUid keep working: the execution layer
falls back to refId at DAG-build time, and the UI backfills the UID on
the next save.
Dependency gates
Each dependency is a pair { itemUid, condition }. The condition decides
when the dependent item becomes eligible to run.
| Gate | Dependent runs when the predecessor … | Typical use case |
|---|---|---|
on_success |
finished with status passed. |
Main happy path. Default. |
on_fail |
finished with status failed. |
Cleanup / rollback branches. |
always |
reached any terminal status (passed/failed/skipped/cancelled). | Teardown / resource cleanup. |
on_skipped |
was skipped (e.g. its own gate was not met). | Recovery branches. |
An empty / missing condition is treated as on_success — the safe
historical default. When a gate is not met, the dependent item is marked
skipped with skipReason: gate_not_met.
Example — three items, two of which share a cleanup step:
{
"name": "API regression",
"executionMode": "dag",
"items": [
{
"itemUid": "11111111-1111-1111-1111-111111111111",
"type": "functional",
"refId": "aaaa...",
"order": 1,
"name": "Auth smoke"
},
{
"itemUid": "22222222-2222-2222-2222-222222222222",
"type": "fuzz",
"refId": "bbbb...",
"order": 2,
"name": "Auth fuzz",
"gates": [
{ "itemUid": "11111111-1111-1111-1111-111111111111", "condition": "on_success" }
]
},
{
"itemUid": "33333333-3333-3333-3333-333333333333",
"type": "functional",
"refId": "cccc...",
"order": 3,
"name": "Teardown",
"gates": [
{ "itemUid": "11111111-1111-1111-1111-111111111111", "condition": "always" },
{ "itemUid": "22222222-2222-2222-2222-222222222222", "condition": "always" }
]
}
]
}
The teardown step runs regardless of whether the smoke or the fuzz
succeeded, failed, or was skipped.
Legacy dependsOn
Plans saved before the redesign used dependsOn: [UUID, ...] on each item.
The server still reads this field: each entry is expanded into a typed
{ itemUid, condition: on_success } edge at runtime. New code should write
the typed gates[] shape instead — it is strictly more expressive.
Retry policy
Any item may carry a retry object that controls failure-retry behaviour.
"retry": {
"maxAttempts": 3,
"backoffMs": 5000,
"onlyOn": ["timeout", "connection reset"]
}
| Field | Range | Meaning |
|---|---|---|
maxAttempts |
1–10 | Total attempts including the first (so 3 = initial + up to 2 retries). 1 disables retry. |
backoffMs |
0–300000 ms | Base linear backoff. Actual delay is backoffMs × attemptIndex. 0 retries immediately. |
onlyOn |
array of strings | Case-insensitive substring matches on the error text. Empty = retry on any error. |
A retried attempt sees a fresh per-attempt context: the TimeoutMs applies
per attempt, not per item. On success the remaining attempts are
skipped. On context cancellation (run cancel or scheduler stop), the
currently-running attempt is aborted and no further retries start.
Per-item timeout
timeoutMs bounds each attempt with context.WithTimeout. When the
deadline fires:
- the attempt is treated as
failedwitherror = "timeout"; timeoutHitis set totrueon the item state;- if retries remain and
onlyOnallows it, the next attempt starts after
the configured backoff.
Valid range: 0–14 400 000 ms (4 hours). 0 means “no per-item deadline
override” (the orchestrator-wide default still applies). Anything longer
should be split into separate scheduled runs.
Critical items
Mark an item critical: true when its failure must stop the whole run.
Semantics:
- If the item finishes with status
failed, the orchestrator cancels
every other in-flight or pending item. - Cancelled siblings carry
status: cancelled,
cancelCause: dependency_critical_failure. - The run itself finishes with
status: failed.
This is different from simply depending on the item: siblings that have
already finished are left alone; only not-yet-done work is cascaded.
Use critical for preconditions (setup, auth, migration) where running
the rest of the plan after a failure is pointless or unsafe.
ContinueOnError
continueOnError: true does the opposite: it says “even if this item
fails, downstream on_success gates should behave as if it had passed”.
The actual run status is still counted as failed in the run’s
failedItems tally, but downstream work continues instead of skipping.
Useful for flaky items whose failure you want logged but not enforced.
on_fail and on_skipped gates are unaffected — they still inspect the
real status.
Live run view
The Run Detail page (/ui/test-plans/<id>/runs/<runId>) opens six tabs:
| Tab | What you see |
|---|---|
| Overview | Run envelope (status, started/completed, duration, triggered-by), item totals, status chip. |
| Graph | Cytoscape DAG with live status colouring. Click a node to open the side panel. |
| Timeline | Gantt-style bars: rows = items, X-axis = wall-clock. Dependencies drawn as dashed lines. Best view for parallel runs. |
| Items | Flat table: order, name, type, status, attempts, duration. |
| Logs | Live tail per item. Filter by item + level + regex. Auto-scroll with a pause toggle. |
| Artifacts | Attachment list (allure.zip, har.json, screenshots, JUnit). Inline preview for small text / JSON. |
| Report | Allure-rendered step tree + download buttons for every export format. |
| Compare | Pick a second run and see the per-item diff (available on any completed run). |
Updates stream over SSE. If the browser disconnects, reconnect is
transparent — the stream resumes from the last seen event ID via the
Last-Event-ID header, so you never lose updates.
Export formats
Every run produces reports in six formats. Choose the one your downstream
tooling speaks — the server generates each on demand.
| Endpoint | MIME type | Typical use |
|---|---|---|
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report |
application/json |
Allure JSON summary. |
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.zip |
application/zip |
Allure archive (result-*.json + attachments). |
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.junit.xml |
application/xml |
JUnit XML for Jenkins / GitLab / GitHub Actions. |
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.md |
text/markdown |
Slack / email / wiki summary. |
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.html |
text/html |
Self-contained HTML, open in any browser, Save-as-PDF. |
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.unified.json |
application/json |
Native Mockarty envelope (plan + counts + per-item state). |
All formats are byte-deterministic given the same run inputs — useful for
snapshot tests. The Allure endpoints honour If-None-Match with a strong
ETag; the others regenerate on every request (cheap).
CLI
mockarty-cli testplan report <runID> --plan plan-abc --format allure -o report.json
mockarty-cli testplan report <runID> --plan plan-abc --format zip -o report.zip
mockarty-cli testplan report <runID> --plan plan-abc --format junit -o report.junit.xml
mockarty-cli testplan report <runID> --plan plan-abc --format markdown -o report.md
mockarty-cli testplan report <runID> --plan plan-abc --format html -o report.html
mockarty-cli testplan report <runID> --plan plan-abc --format unified -o report.unified.json
SDK
| Go | Python | Java |
|---|---|---|
GetRunReport |
get_run_report |
getRunReport |
GetRunReportZIP |
get_run_report_zip |
getRunReportZip |
GetRunReportJUnit |
get_run_report_junit |
getRunReportJUnit |
GetRunReportMarkdown |
get_run_report_markdown |
getRunReportMarkdown |
GetRunReportHTML |
get_run_report_html |
getRunReportHTML |
GetRunReportUnified |
get_run_report_unified |
getRunReportUnified |
Compare runs
Diff any two runs side by side. Both runs must live in the caller’s
namespace. Comparing runs of different plans is allowed — the response
sets summary.differentPlans: true so the UI can show a banner.
GET /api/v1/test-plans/runs/compare?run_a=<runID>&run_b=<runID>
Pass the older run as run_a and the newer run as run_b so regression
and improvement signs read intuitively.
The classifier labels each item as pass_to_fail, fail_to_pass,
skipped_to_ran, ran_to_skipped, fail_to_fail, pass_to_pass,
added, removed, or unchanged. A pass_to_pass item is additionally
flagged durationWorsened when the new run took at least 20% longer
than the baseline and the absolute delta exceeds 250 ms. Small
noise on fast items is therefore ignored.
CLI
mockarty-cli testplan compare-runs <runA> <runB> # table, hides unchanged
mockarty-cli testplan compare-runs <runA> <runB> --diff-only=false # show unchanged too
mockarty-cli testplan compare-runs <runA> <runB> -o json | jq .summary
Exit code is 1 when at least one regression is detected — good for a CI
gate.
Entity search
The in-UI pickers (Add Item, dependency selector, aggregate-report run
selector, schedule target) all call a single search endpoint:
GET /api/v1/entity-search?type=<kind>&q=<name>&namespace=<ns>&limit=20&offset=0
Supported type values: mock, test_plan, perf_config,
fuzz_config, chaos_experiment, contract_pact. q is a
case-insensitive substring match on the entity name; response carries
{ items: [{id, type, name, namespace, createdAt, numericId}], total }.
See the dedicated Entity Search page for the full shape,
SDK + CLI examples, and MCP tool details.
Distributed runner path (optional)
By default the admin node executes every Test Plan item in-process. When
you want to offload item execution to a dedicated mockarty-runner
worker — for isolation, network placement, or simply scale — enable the
claim queue:
- Admin node: set
MOCKARTY_RUNNER_TESTPLAN_ENABLED=true. The
orchestrator now publishes items to an in-memory claim queue instead
of executing them locally. - Runner node: set
RUNNER_TESTPLAN_ENABLED=trueand
COORDINATOR_URL=https://<admin>. Optional:
RUNNER_TESTPLAN_CONCURRENCY(default4) bounds how many items the
runner claims in parallel. - Per-item targeting: populate
runnerLabels: ["gpu", "staging"]
on items you want a specific runner pool to pick up. Empty labels
mean “any runner whose namespace matches”. For more expressive
targeting (OR, NOT, regex), userunnerLabelExprinstead — see
Runner Labels and Targeting. The two fields are
AND-combined when both are set; an expression-bearing item bypasses
the in-memory claim queue and dispatches through the coordinator so
the runner-side DSL evaluator can honour the AST.
The runner long-polls three endpoints:
| Method | Path | Purpose |
|---|---|---|
| POST | /api/v1/runner/testplan/claim |
Claim the next matching item (long-poll, max 30s). |
| POST | /api/v1/runner/testplan/report |
Report a terminal status + summary + artifacts. |
| POST | /api/v1/runner/testplan/heartbeat |
Keep the claim alive while the item runs. |
When the admin node is restarted or the claim’s heartbeat deadline
elapses, unclaimed work is requeued automatically — the runner path has
at-least-once semantics, and runners must make their own attempts
idempotent.
If MOCKARTY_RUNNER_TESTPLAN_ENABLED is unset on the admin node, all
three endpoints respond with 503 Service Unavailable and the runner
gracefully falls back to the classic task protocol — existing runners
never break.
SDK / CLI / MCP
Every endpoint on this page is available through:
- Go / Python / Java SDKs — typed methods for create, run, report,
compare, schedule, and webhook CRUD. See SDK Guide. - CLI —
mockarty-cli testplan <create|run|report|compare-runs| schedule|webhook|stream>. See CLI User Guide. - MCP tools — agents call
list_test_plans,run_test_plan,
get_test_plan_run,compare_test_plan_runs,search_entities, and
the exports (get_test_plan_run_report,…_junit,…_markdown,
…_html,…_unified). See AI Features.
Backward compatibility
- Prefer
executionMode(fifo/parallel/dag) for new
integrations. The olderschedulevalue is still accepted and kept
consistent, so existing integrations keep working unchanged. - Plans created before the redesign used
dependsOn: [UUID, ...]. The
orchestrator still reads this list and converts each entry into a
typedon_successgate at runtime. Re-saving the plan in the new UI
materialises the typedgates[]shape. items[].itemUidis generated on first save. Legacy plans that never
received one keep working — the orchestrator falls back torefIdfor
dependency resolution until the next save.
Troubleshooting
/api/v1/runner/testplan/*returns503— claim queue is
disabled. SetMOCKARTY_RUNNER_TESTPLAN_ENABLED=trueon the admin
node and restart.- Item stays
pendingforever — with runner labels set, check that
at least onemockarty-runnercarries a matching label superset. The
admin’s/metricsexposestestplan_claim_queue_depthfor a
dashboard; the runner’s/metricsexposestestplan_runner_claims_total. - Retries never kick in — verify
retry.maxAttempts > 1. If
onlyOnis set, make sure the error text your executor returns
actually contains one of the configured substrings (the match is
case-insensitive). - Critical cascade didn’t fire — only items that were still pending
or running when the critical item failed are cancelled. Items that had
already finished keep their final state. - Live stream disconnects drop events — the browser sends
Last-Event-IDon reconnect and the server resumes. If you hit a
corporate proxy that strips SSE headers, fall back to polling
GET /runs/:idon a short interval.
Per-attempt timeline (Gantt stacking)
When an item has a retry policy, the Timeline tab renders one sub-bar per
attempt stacked left-to-right instead of a single aggregate bar. Each
sub-bar is coloured by that attempt’s terminal status (passed / failed /
timeout) so transient failures followed by a successful retry are visible
at a glance.
What the UI shows:
- Sub-bar label — the attempt index (
1,2,3, …). Width is
proportional to that attempt’s execution window. - Tooltip — status, attempt index / total, duration, optional
timeout marker, runner id, and per-attempt error text (truncated to
200 chars). - Timeout glyph — sub-bars whose attempt was terminated by the
per-item timeout render with a dashed danger border. - Running attempts — in-flight attempts pulse until the finish event
arrives via SSE. No page refresh required.
Payload shape: each ItemState row returned by GET /test-plan-runs/:id
now carries an attemptLog array:
{
"itemUid": "…",
"status": "passed",
"attempts": 3,
"attemptLog": [
{"startedAt": "2026-04-20T10:00:00Z", "completedAt": "2026-04-20T10:00:02Z",
"durationMs": 2000, "status": "failed", "error": "transient"},
{"startedAt": "2026-04-20T10:00:03Z", "completedAt": "2026-04-20T10:00:04Z",
"durationMs": 1000, "status": "failed", "error": "transient"},
{"startedAt": "2026-04-20T10:00:05Z", "completedAt": "2026-04-20T10:00:06Z",
"durationMs": 1000, "status": "passed"}
]
}
Runs created before this release have attemptLog empty — the UI then
falls back to the legacy single-bar render annotated with the attempt
count (×3).
Timeline correlation (overlay runs)
Below the main Timeline you can overlay one or more previous runs of the
same plan onto a shared time axis. Use the Add run picker to pull
another run’s itemsState and render it as a separate lane beneath the
primary timeline. Each lane reuses the same sub-bar logic (stacked
attempts, timeout glyph, status colouring) so drift between runs — a
step that used to finish in 2 s now taking 8 s, a retry that didn’t fire
in this run — is visible without switching tabs.
Correlation is client-side only: picking a run fires a single
GET /api/v1/test-plans/runs/:id and renders locally. No server-side
join, no extra load on the orchestrator.
Aggregate reports across runs
An aggregate report folds two or more existing test runs into a single
release-ready report. Use it when one logical release gate spans several
independent executions — a fuzz campaign, a chaos experiment, and a
functional regression pass — and you want one pass/fail summary and a single
report to share with stakeholders.
The report is stateless: nothing new is stored. You pass the run IDs on
each call and the server recomputes the report on the fly. The source runs
keep their own history, reports, and retention — untouched.
Building an aggregate report
# CLI — HTML is self-contained and print-to-PDF friendly
mockarty-cli testplan aggregate-report <run-id-1> <run-id-2> [<run-id-3> ...] \
--name "Release 1.4 gate" --format html -o release.html
// Go SDK
data, err := client.TestRuns().AggregateRunsReport(ctx,
mockarty.AggregateRunsReportRequest{
Name: "Release 1.4 gate",
RunIDs: []string{runAID, runBID, runCID},
},
mockarty.AggregateReportFormatHTML)
All run IDs must live in the caller’s namespace (admins and support can span
namespaces). To report on a different set of runs, simply call again with the
new IDs — there is no entity to create, update, or delete.
Formats
| Format | MIME | Use case |
|---|---|---|
unified |
application/json |
Programmatic consumers (SDK / CLI / CI) |
markdown |
text/markdown |
Slack / wiki paste, human review |
html |
text/html |
Release report — self-contained, print-to-PDF |
junit |
application/xml |
CI ingest (Jenkins / GitLab / etc.) |
mockarty-cli testplan aggregate-report r1 r2 --format junit -o ci.xml
mockarty-cli testplan aggregate-report r1 r2 --format markdown -o out.md
HTML output inlines its CSS and charts, so it renders in an air-gapped
browser and saves cleanly as PDF via the browser print dialog.
When to use an aggregate report vs. a Test Plan
- Test Plan — you decide ahead of time that a group of items runs
together, with explicit dependencies, retries, gates, and a shared
report. Good for release pipelines and nightlies. - Aggregate report — you want to summarise after the fact. The runs
already exist, possibly from different CI jobs or on-demand executions, and
you just need one report to share. No execution orchestration, no stored
entity — purely a recomputed view.