Docs Advanced Test Plans

Advanced Test Plans

Companion guide to Test Plans. This page covers the
execution-layer features:

  • typed dependency gates (on_success / on_fail / always / on_skipped)
  • per-item retry, timeout, critical and continueOnError flags
  • stable itemUid dependency keys
  • the live run view (Graph / Timeline / Items / Logs / Artifacts / Report)
  • six report formats and run-to-run comparison
  • the optional distributed runner path for plan items
  • the entity-search backend that powers every in-UI picker

About URLs in examples: we use http://localhost:5770 as the default
Mockarty address. Replace it with your own installation URL if different.
See Tips & Useful Features.

Related pages: Test Plans ·
Test Plans API Cookbook ·
Test Plans in CI/CD ·
Entity Search ·
SDK Guide

Execution modes

Every plan is saved with one of three modes (executionMode field).

Mode When to use How it runs
fifo Legacy sequential plans; one step at a time. Items run in ascending order; a failure short-circuits remaining items (unless overridden).
parallel Independent items that can all start at once. All items dispatched concurrently; no dependency edges honoured.
dag Typed dependencies with gates (default for new plans). Items run in dependency order; each gate decides when its item becomes eligible.

If you omit executionMode, the server auto-detects dag when any item
carries typed gates, otherwise fifo. For backward compatibility the
older schedule value (parallel / dag) is still accepted.

Items and dependencies

ItemUID — stable reference key

Every item in a plan carries an itemUid (UUID). It is generated server-side
when the plan is first saved and stays stable across edits. Dependency edges
point at itemUid, not at the external entity UUID (refId), so:

  • you can add the same collection / config twice (e.g. smoke + regression +
    cleanup) without ambiguity;
  • swapping the underlying entity (via refId edit) does not break the plan
    graph;
  • renaming or deleting an entity only affects that one item, not siblings.

Legacy plans that pre-date itemUid keep working: the execution layer
falls back to refId at DAG-build time, and the UI backfills the UID on
the next save.

Dependency gates

Each dependency is a pair { itemUid, condition }. The condition decides
when the dependent item becomes eligible to run.

Gate Dependent runs when the predecessor … Typical use case
on_success finished with status passed. Main happy path. Default.
on_fail finished with status failed. Cleanup / rollback branches.
always reached any terminal status (passed/failed/skipped/cancelled). Teardown / resource cleanup.
on_skipped was skipped (e.g. its own gate was not met). Recovery branches.

An empty / missing condition is treated as on_success — the safe
historical default. When a gate is not met, the dependent item is marked
skipped with skipReason: gate_not_met.

Example — three items, two of which share a cleanup step:

{
  "name": "API regression",
  "executionMode": "dag",
  "items": [
    {
      "itemUid": "11111111-1111-1111-1111-111111111111",
      "type": "functional",
      "refId": "aaaa...",
      "order": 1,
      "name": "Auth smoke"
    },
    {
      "itemUid": "22222222-2222-2222-2222-222222222222",
      "type": "fuzz",
      "refId": "bbbb...",
      "order": 2,
      "name": "Auth fuzz",
      "gates": [
        { "itemUid": "11111111-1111-1111-1111-111111111111", "condition": "on_success" }
      ]
    },
    {
      "itemUid": "33333333-3333-3333-3333-333333333333",
      "type": "functional",
      "refId": "cccc...",
      "order": 3,
      "name": "Teardown",
      "gates": [
        { "itemUid": "11111111-1111-1111-1111-111111111111", "condition": "always" },
        { "itemUid": "22222222-2222-2222-2222-222222222222", "condition": "always" }
      ]
    }
  ]
}

The teardown step runs regardless of whether the smoke or the fuzz
succeeded, failed, or was skipped.

Legacy dependsOn

Plans saved before the redesign used dependsOn: [UUID, ...] on each item.
The server still reads this field: each entry is expanded into a typed
{ itemUid, condition: on_success } edge at runtime. New code should write
the typed gates[] shape instead — it is strictly more expressive.

Retry policy

Any item may carry a retry object that controls failure-retry behaviour.

"retry": {
  "maxAttempts": 3,
  "backoffMs": 5000,
  "onlyOn": ["timeout", "connection reset"]
}
Field Range Meaning
maxAttempts 1–10 Total attempts including the first (so 3 = initial + up to 2 retries). 1 disables retry.
backoffMs 0–300000 ms Base linear backoff. Actual delay is backoffMs × attemptIndex. 0 retries immediately.
onlyOn array of strings Case-insensitive substring matches on the error text. Empty = retry on any error.

A retried attempt sees a fresh per-attempt context: the TimeoutMs applies
per attempt, not per item. On success the remaining attempts are
skipped. On context cancellation (run cancel or scheduler stop), the
currently-running attempt is aborted and no further retries start.

Per-item timeout

timeoutMs bounds each attempt with context.WithTimeout. When the
deadline fires:

  • the attempt is treated as failed with error = "timeout";
  • timeoutHit is set to true on the item state;
  • if retries remain and onlyOn allows it, the next attempt starts after
    the configured backoff.

Valid range: 0–14 400 000 ms (4 hours). 0 means “no per-item deadline
override” (the orchestrator-wide default still applies). Anything longer
should be split into separate scheduled runs.

Critical items

Mark an item critical: true when its failure must stop the whole run.
Semantics:

  • If the item finishes with status failed, the orchestrator cancels
    every other in-flight or pending item.
  • Cancelled siblings carry status: cancelled,
    cancelCause: dependency_critical_failure.
  • The run itself finishes with status: failed.

This is different from simply depending on the item: siblings that have
already finished are left alone; only not-yet-done work is cascaded.

Use critical for preconditions (setup, auth, migration) where running
the rest of the plan after a failure is pointless or unsafe.

ContinueOnError

continueOnError: true does the opposite: it says “even if this item
fails, downstream on_success gates should behave as if it had passed”.
The actual run status is still counted as failed in the run’s
failedItems tally, but downstream work continues instead of skipping.

Useful for flaky items whose failure you want logged but not enforced.

on_fail and on_skipped gates are unaffected — they still inspect the
real status.

Live run view

The Run Detail page (/ui/test-plans/<id>/runs/<runId>) opens six tabs:

Tab What you see
Overview Run envelope (status, started/completed, duration, triggered-by), item totals, status chip.
Graph Cytoscape DAG with live status colouring. Click a node to open the side panel.
Timeline Gantt-style bars: rows = items, X-axis = wall-clock. Dependencies drawn as dashed lines. Best view for parallel runs.
Items Flat table: order, name, type, status, attempts, duration.
Logs Live tail per item. Filter by item + level + regex. Auto-scroll with a pause toggle.
Artifacts Attachment list (allure.zip, har.json, screenshots, JUnit). Inline preview for small text / JSON.
Report Allure-rendered step tree + download buttons for every export format.
Compare Pick a second run and see the per-item diff (available on any completed run).

Updates stream over SSE. If the browser disconnects, reconnect is
transparent — the stream resumes from the last seen event ID via the
Last-Event-ID header, so you never lose updates.

Export formats

Every run produces reports in six formats. Choose the one your downstream
tooling speaks — the server generates each on demand.

Endpoint MIME type Typical use
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report application/json Allure JSON summary.
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.zip application/zip Allure archive (result-*.json + attachments).
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.junit.xml application/xml JUnit XML for Jenkins / GitLab / GitHub Actions.
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.md text/markdown Slack / email / wiki summary.
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.html text/html Self-contained HTML, open in any browser, Save-as-PDF.
GET /api/v1/namespaces/:ns/test-plans/:idOrNumericID/runs/:runID/report.unified.json application/json Native Mockarty envelope (plan + counts + per-item state).

All formats are byte-deterministic given the same run inputs — useful for
snapshot tests. The Allure endpoints honour If-None-Match with a strong
ETag; the others regenerate on every request (cheap).

CLI

mockarty-cli testplan report <runID> --plan plan-abc --format allure   -o report.json
mockarty-cli testplan report <runID> --plan plan-abc --format zip      -o report.zip
mockarty-cli testplan report <runID> --plan plan-abc --format junit    -o report.junit.xml
mockarty-cli testplan report <runID> --plan plan-abc --format markdown -o report.md
mockarty-cli testplan report <runID> --plan plan-abc --format html     -o report.html
mockarty-cli testplan report <runID> --plan plan-abc --format unified  -o report.unified.json

SDK

Go Python Java
GetRunReport get_run_report getRunReport
GetRunReportZIP get_run_report_zip getRunReportZip
GetRunReportJUnit get_run_report_junit getRunReportJUnit
GetRunReportMarkdown get_run_report_markdown getRunReportMarkdown
GetRunReportHTML get_run_report_html getRunReportHTML
GetRunReportUnified get_run_report_unified getRunReportUnified

Compare runs

Diff any two runs side by side. Both runs must live in the caller’s
namespace. Comparing runs of different plans is allowed — the response
sets summary.differentPlans: true so the UI can show a banner.

GET /api/v1/test-plans/runs/compare?run_a=<runID>&run_b=<runID>

Pass the older run as run_a and the newer run as run_b so regression
and improvement signs read intuitively.

The classifier labels each item as pass_to_fail, fail_to_pass,
skipped_to_ran, ran_to_skipped, fail_to_fail, pass_to_pass,
added, removed, or unchanged. A pass_to_pass item is additionally
flagged durationWorsened when the new run took at least 20% longer
than the baseline and the absolute delta exceeds 250 ms. Small
noise on fast items is therefore ignored.

CLI

mockarty-cli testplan compare-runs <runA> <runB>                     # table, hides unchanged
mockarty-cli testplan compare-runs <runA> <runB> --diff-only=false   # show unchanged too
mockarty-cli testplan compare-runs <runA> <runB> -o json | jq .summary

Exit code is 1 when at least one regression is detected — good for a CI
gate.

The in-UI pickers (Add Item, dependency selector, aggregate-report run
selector, schedule target) all call a single search endpoint:

GET /api/v1/entity-search?type=<kind>&q=<name>&namespace=<ns>&limit=20&offset=0

Supported type values: mock, test_plan, perf_config,
fuzz_config, chaos_experiment, contract_pact. q is a
case-insensitive substring match on the entity name; response carries
{ items: [{id, type, name, namespace, createdAt, numericId}], total }.

See the dedicated Entity Search page for the full shape,
SDK + CLI examples, and MCP tool details.

Distributed runner path (optional)

By default the admin node executes every Test Plan item in-process. When
you want to offload item execution to a dedicated mockarty-runner
worker — for isolation, network placement, or simply scale — enable the
claim queue:

  1. Admin node: set MOCKARTY_RUNNER_TESTPLAN_ENABLED=true. The
    orchestrator now publishes items to an in-memory claim queue instead
    of executing them locally.
  2. Runner node: set RUNNER_TESTPLAN_ENABLED=true and
    COORDINATOR_URL=https://<admin>. Optional:
    RUNNER_TESTPLAN_CONCURRENCY (default 4) bounds how many items the
    runner claims in parallel.
  3. Per-item targeting: populate runnerLabels: ["gpu", "staging"]
    on items you want a specific runner pool to pick up. Empty labels
    mean “any runner whose namespace matches”. For more expressive
    targeting (OR, NOT, regex), use runnerLabelExpr instead — see
    Runner Labels and Targeting. The two fields are
    AND-combined when both are set; an expression-bearing item bypasses
    the in-memory claim queue and dispatches through the coordinator so
    the runner-side DSL evaluator can honour the AST.

The runner long-polls three endpoints:

Method Path Purpose
POST /api/v1/runner/testplan/claim Claim the next matching item (long-poll, max 30s).
POST /api/v1/runner/testplan/report Report a terminal status + summary + artifacts.
POST /api/v1/runner/testplan/heartbeat Keep the claim alive while the item runs.

When the admin node is restarted or the claim’s heartbeat deadline
elapses, unclaimed work is requeued automatically — the runner path has
at-least-once semantics, and runners must make their own attempts
idempotent.

If MOCKARTY_RUNNER_TESTPLAN_ENABLED is unset on the admin node, all
three endpoints respond with 503 Service Unavailable and the runner
gracefully falls back to the classic task protocol — existing runners
never break.

SDK / CLI / MCP

Every endpoint on this page is available through:

  • Go / Python / Java SDKs — typed methods for create, run, report,
    compare, schedule, and webhook CRUD. See SDK Guide.
  • CLI — mockarty-cli testplan <create|run|report|compare-runs| schedule|webhook|stream>. See CLI User Guide.
  • MCP tools — agents call list_test_plans, run_test_plan,
    get_test_plan_run, compare_test_plan_runs, search_entities, and
    the exports (get_test_plan_run_report, …_junit, …_markdown,
    …_html, …_unified). See AI Features.

Backward compatibility

  • Prefer executionMode (fifo / parallel / dag) for new
    integrations. The older schedule value is still accepted and kept
    consistent, so existing integrations keep working unchanged.
  • Plans created before the redesign used dependsOn: [UUID, ...]. The
    orchestrator still reads this list and converts each entry into a
    typed on_success gate at runtime. Re-saving the plan in the new UI
    materialises the typed gates[] shape.
  • items[].itemUid is generated on first save. Legacy plans that never
    received one keep working — the orchestrator falls back to refId for
    dependency resolution until the next save.

Troubleshooting

  • /api/v1/runner/testplan/* returns 503 — claim queue is
    disabled. Set MOCKARTY_RUNNER_TESTPLAN_ENABLED=true on the admin
    node and restart.
  • Item stays pending forever — with runner labels set, check that
    at least one mockarty-runner carries a matching label superset. The
    admin’s /metrics exposes testplan_claim_queue_depth for a
    dashboard; the runner’s /metrics exposes testplan_runner_claims_total.
  • Retries never kick in — verify retry.maxAttempts > 1. If
    onlyOn is set, make sure the error text your executor returns
    actually contains one of the configured substrings (the match is
    case-insensitive).
  • Critical cascade didn’t fire — only items that were still pending
    or running when the critical item failed are cancelled. Items that had
    already finished keep their final state.
  • Live stream disconnects drop events — the browser sends
    Last-Event-ID on reconnect and the server resumes. If you hit a
    corporate proxy that strips SSE headers, fall back to polling
    GET /runs/:id on a short interval.

Per-attempt timeline (Gantt stacking)

When an item has a retry policy, the Timeline tab renders one sub-bar per
attempt stacked left-to-right instead of a single aggregate bar. Each
sub-bar is coloured by that attempt’s terminal status (passed / failed /
timeout) so transient failures followed by a successful retry are visible
at a glance.

What the UI shows:

  • Sub-bar label — the attempt index (1, 2, 3, …). Width is
    proportional to that attempt’s execution window.
  • Tooltip — status, attempt index / total, duration, optional
    timeout marker, runner id, and per-attempt error text (truncated to
    200 chars).
  • Timeout glyph — sub-bars whose attempt was terminated by the
    per-item timeout render with a dashed danger border.
  • Running attempts — in-flight attempts pulse until the finish event
    arrives via SSE. No page refresh required.

Payload shape: each ItemState row returned by GET /test-plan-runs/:id
now carries an attemptLog array:

{
  "itemUid": "…",
  "status": "passed",
  "attempts": 3,
  "attemptLog": [
    {"startedAt": "2026-04-20T10:00:00Z", "completedAt": "2026-04-20T10:00:02Z",
     "durationMs": 2000, "status": "failed", "error": "transient"},
    {"startedAt": "2026-04-20T10:00:03Z", "completedAt": "2026-04-20T10:00:04Z",
     "durationMs": 1000, "status": "failed", "error": "transient"},
    {"startedAt": "2026-04-20T10:00:05Z", "completedAt": "2026-04-20T10:00:06Z",
     "durationMs": 1000, "status": "passed"}
  ]
}

Runs created before this release have attemptLog empty — the UI then
falls back to the legacy single-bar render annotated with the attempt
count (×3).

Timeline correlation (overlay runs)

Below the main Timeline you can overlay one or more previous runs of the
same plan onto a shared time axis. Use the Add run picker to pull
another run’s itemsState and render it as a separate lane beneath the
primary timeline. Each lane reuses the same sub-bar logic (stacked
attempts, timeout glyph, status colouring) so drift between runs — a
step that used to finish in 2 s now taking 8 s, a retry that didn’t fire
in this run — is visible without switching tabs.

Correlation is client-side only: picking a run fires a single
GET /api/v1/test-plans/runs/:id and renders locally. No server-side
join, no extra load on the orchestrator.

Aggregate reports across runs

An aggregate report folds two or more existing test runs into a single
release-ready report. Use it when one logical release gate spans several
independent executions — a fuzz campaign, a chaos experiment, and a
functional regression pass — and you want one pass/fail summary and a single
report to share with stakeholders.

The report is stateless: nothing new is stored. You pass the run IDs on
each call and the server recomputes the report on the fly. The source runs
keep their own history, reports, and retention — untouched.

Building an aggregate report

# CLI — HTML is self-contained and print-to-PDF friendly
mockarty-cli testplan aggregate-report <run-id-1> <run-id-2> [<run-id-3> ...] \
    --name "Release 1.4 gate" --format html -o release.html
// Go SDK
data, err := client.TestRuns().AggregateRunsReport(ctx,
    mockarty.AggregateRunsReportRequest{
        Name:   "Release 1.4 gate",
        RunIDs: []string{runAID, runBID, runCID},
    },
    mockarty.AggregateReportFormatHTML)

All run IDs must live in the caller’s namespace (admins and support can span
namespaces). To report on a different set of runs, simply call again with the
new IDs — there is no entity to create, update, or delete.

Formats

Format MIME Use case
unified application/json Programmatic consumers (SDK / CLI / CI)
markdown text/markdown Slack / wiki paste, human review
html text/html Release report — self-contained, print-to-PDF
junit application/xml CI ingest (Jenkins / GitLab / etc.)
mockarty-cli testplan aggregate-report r1 r2 --format junit -o ci.xml
mockarty-cli testplan aggregate-report r1 r2 --format markdown -o out.md

HTML output inlines its CSS and charts, so it renders in an air-gapped
browser and saves cleanly as PDF via the browser print dialog.

When to use an aggregate report vs. a Test Plan

  • Test Plan — you decide ahead of time that a group of items runs
    together, with explicit dependencies, retries, gates, and a shared
    report. Good for release pipelines and nightlies.
  • Aggregate report — you want to summarise after the fact. The runs
    already exist, possibly from different CI jobs or on-demand executions, and
    you just need one report to share. No execution orchestration, no stored
    entity — purely a recomputed view.