Docs Ephemeral Runners (CI / Kubernetes)

Ephemeral Runners

A Mockarty ephemeral runner is a runner process whose lifetime is bounded by a single CI job (or a fixed idle window). It registers, picks up its work, then exits — and the coordinator removes the runner row immediately instead of leaving it in the offline state for the next reconciliation pass.

Use ephemeral runners when:

  • The runner lives inside a CI container (GitLab job, GitHub Actions container job, Jenkins agent) that disappears when the pipeline finishes.
  • You spin runners on demand inside Kubernetes (HPA, Job, KEDA), and you do not want the admin UI to fill up with stale runner rows after pods exit.
  • You want a CI step to time out cleanly after N seconds of idle so the container does not linger on the runner — let the runner exit itself when it has no work.

Long-lived runners (a server you start once and forget) keep their default behaviour — no opt-in needed.

Quick start

Pass EPHEMERAL=true when launching the runner. Optionally combine with IDLE_TIMEOUT so the process exits when it has nothing to do.

Before using the examples, set the shell or CI variable MOCKARTY_RUNNER_IMAGE to the exact immutable repository@sha256:... runner image from your verified release set. It is an example variable, not a Mockarty server setting. Mirror the image into the private registry first in an air-gapped installation; do not substitute a mutable latest tag.

Docker run

docker run --rm \
  -e API_TOKEN=mki_your_integration_token \
  -e COORDINATOR_ADDR=mockarty.example.com:5773 \
  -e RUNNER_NAME=ci-runner-${CI_JOB_ID} \
  -e LABELS=ci,gitlab,linux \
  -e EPHEMERAL=true \
  -e IDLE_TIMEOUT=120s \
  "$MOCKARTY_RUNNER_IMAGE"

When the runner finishes its work (no active tasks for 120 seconds) it deregisters cleanly and the container exits. The runner row is removed from the runner list on the next refresh.

Docker Compose (one-shot CI helper)

services:
  runner:
    image: ${MOCKARTY_RUNNER_IMAGE:?set an immutable repository@sha256 digest}
    restart: "no"
    environment:
      API_TOKEN: ${MOCKARTY_INTEGRATION_TOKEN}
      COORDINATOR_ADDR: ${MOCKARTY_COORDINATOR_ADDR}
      RUNNER_NAME: ci-${CI_PIPELINE_ID}-${CI_JOB_ID}
      LABELS: ci,gitlab,${CI_RUNNER_TAGS}
      EPHEMERAL: "true"
      IDLE_TIMEOUT: 90s

GitLab .gitlab-ci.yml

Mirror and pin the Docker CLI and DinD service as part of the same verified CI release set; replace both digest placeholders below.

performance-test:
  image: <PRIVATE_REGISTRY>/docker-cli@sha256:<DOCKER_CLI_IMAGE_DIGEST>
  services:
    - name: <PRIVATE_REGISTRY>/docker-dind@sha256:<DOCKER_DIND_IMAGE_DIGEST>
  variables:
    MOCKARTY_INTEGRATION_TOKEN: $MOCKARTY_RUNNER_TOKEN
  script:
    - |
      docker run --rm \
        -e API_TOKEN=$MOCKARTY_INTEGRATION_TOKEN \
        -e COORDINATOR_ADDR=mockarty.example.com:5773 \
        -e RUNNER_NAME=gitlab-$CI_PROJECT_NAME-$CI_JOB_ID \
        -e LABELS=ci,gitlab \
        -e EPHEMERAL=true \
        -e IDLE_TIMEOUT=300s \
        "$MOCKARTY_RUNNER_IMAGE"

Kubernetes Job

apiVersion: batch/v1
kind: Job
metadata:
  generateName: mockarty-runner-
spec:
  ttlSecondsAfterFinished: 60
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: runner
        image: <PRIVATE_REGISTRY>/mockarty-runner@sha256:<RUNNER_IMAGE_DIGEST>
        env:
        - name: API_TOKEN
          valueFrom:
            secretKeyRef: { name: mockarty-runner, key: token }
        - name: COORDINATOR_ADDR
          value: mockarty.example.com:5773
        - name: EPHEMERAL
          value: "true"
        - name: IDLE_TIMEOUT
          value: "180s"
        - name: LABELS
          value: "ci,k8s,linux"

One-task and drain modes (--once / --drain)

Besides the idle window, the general runner (mockarty-runner) has two explicit CI lifecycle modes. Pass them as command-line flags (they win over the environment) or set the matching variable:

Mode Flag Environment (canonical / legacy) What the runner does
One task --once MOCKARTY_RUNNER_ONCE=true / RUNNER_ONCE=true Claims exactly one task, waits for it to finish, deregisters, exits 0.
Drain --drain MOCKARTY_RUNNER_DRAIN=true / RUNNER_DRAIN=true Claims no new tasks at all, waits for anything already running in this process, deregisters, exits 0.
# CI job container that runs exactly one task and disappears:
mockarty-runner --once

# Graceful stop that must not claim a final task:
mockarty-runner --drain

The two modes are mutually exclusive: --once --drain (or both variables set) fails at boot with a configuration error instead of picking one silently.

Combine --once with CLAIM_TOKENS when a CI trigger spawns the runner for one specific dispatch: the runner picks up exactly the token-gated task, runs it, and leaves.

The security runner (mockarty-redteam-runner) supports the same contracts with --once (equivalent to --max-tasks=1) and --drain (refuses new dispatch/lease work and finishes what it holds).

What happens at deregister

When EPHEMERAL=true:

  1. The runner calls DeregisterRunner as part of graceful shutdown (SIGTERM, idle timeout, or the configured DRAIN_TIMEOUT expiring).
  2. The coordinator removes the row from runner_agents (instead of flipping it to status=offline).
  3. The runner list in the admin UI no longer shows the runner — no manual cleanup, no “ghost” rows.

If deregister fails (coordinator unreachable, network blip), the runner retries up to three times with one-second backoff before giving up. Rows orphaned that way are reaped by the leader-only sweeper that runs every 60 seconds on the admin node and deletes ephemeral rows whose last heartbeat is older than three minutes. So even a hard crash (kill -9, OOM, node failure) eventually cleans up.

A long-lived (non-ephemeral) runner was decommissioned and still shows in the list.

  • Evict it explicitly: DELETE /api/v1/runners/<id> force-offlines the row (a later
    reconnect with the same token revives it), and DELETE /api/v1/runners/<id>?purge=true
    removes the row entirely. Requires integrations write permission. A runner that is
    still alive and heartbeating will simply re-appear — purge is for runners that are gone.

Environment variables

Variable Default Description
EPHEMERAL false When true, mark the runner as ephemeral. Deregister becomes a hard delete; the leader-only sweeper catches dropped rows after three minutes.
IDLE_TIMEOUT 0s When > 0 and EPHEMERAL=true, exit gracefully after this much wall-clock time with zero active tasks. Useful for one-shot CI jobs. Accepts Go duration syntax (30s, 2m, 1h) or a bare integer treated as seconds.
MOCKARTY_EPHEMERAL_SWEEP_INTERVAL 60s (admin side) How often the leader-only sweeper checks for stale ephemeral runner rows. Reduce in tests; rarely changed in production.
MOCKARTY_EPHEMERAL_SWEEP_STALE_AFTER 3m (admin side) How old a missed heartbeat must be before the sweeper reaps the row. Increase only if your network’s heartbeat jitter is unusually large.

Required pieces

  • X-Mockarty-Runner-Instance header is mandatory on every request from an ephemeral runner. The Mockarty runner client sets it automatically with a per-process UUID; no operator action needed.
  • Labels stay useful for ephemeral runners — combine EPHEMERAL=true with LABELS=ci,gitlab and your CI pipelines can target dedicated label expressions (e.g. ci & gitlab & !staging) without polluting your long-lived fleet.

Pull mode — runners on team machines (outbound-only)

A runner does not need any inbound network access. By default it runs in pull mode: it dials the admin, long-polls for work, runs the task, and pushes the result back — every connection is runner → admin. This is exactly what you want on a developer’s laptop or a CI worker behind NAT/a firewall: open egress to the admin, nothing inbound. Team members can spin up a runner and take on browser tests, load shards, mobile tests, or security scans without exposing a single port.

The only two settings every runner needs are the admin URL and a runner token:

docker run --rm \
  -e MOCKARTY_ADMIN_URL=https://mockarty.example.com \
  -e MOCKARTY_RUNNER_TOKEN=mki_your_integration_token \
  -e MOCKARTY_RUNNER_NAMESPACES=team-a \
  -e MOCKARTY_RUNNER_LABELS=region=eu,env=dev \
  "$MOCKARTY_RUNNER_IMAGE"

That is enough — pull mode is the default, so no transport flag is required. Mint the token in Admin → Integrations: create an integration of type Test Runner and copy its token.

Common settings (the same names work across runner types):

Variable Default Description
MOCKARTY_ADMIN_URL — Admin base URL the runner dials (https). Outbound only.
MOCKARTY_RUNNER_TOKEN — Integration token; also the runner’s identity.
MOCKARTY_RUNNER_MODE pull pull (outbound long-poll, the default) — set this only if you deliberately want a different transport.
MOCKARTY_RUNNER_INSTANCE per-process Distinguishes multiple runners that share one token.
MOCKARTY_RUNNER_NAMESPACES — Comma-separated namespaces this runner serves.
MOCKARTY_RUNNER_LABELS — Comma-separated key=value selector labels (region=eu,env=dev).
MOCKARTY_RUNNER_MAX_CONCURRENT per-runner Concurrency ceiling.

Each runner type is the same picture with a type-specific image and one or two extra knobs:

  • General runner (API/functional/performance) — the immutable image referenced by MOCKARTY_RUNNER_IMAGE in the example above.
  • Browser UI runner — the browser-runner image; advertises the ui-test capability automatically.
  • Mobile runner — the mobile-runner image; add RUNNER_PLATFORM=android and attach a device/emulator. Its live screen stream uses MOCKARTY_STREAM_MODE=relay (the default), which sends frames through the admin — so the stream is visible even when the runner has no inbound reachability (no WebRTC/TURN required). Set MOCKARTY_STREAM_MODE=webrtc only on a network where peer-to-peer media can connect.
  • Security runners (scanners) — the redteam/exploit-runner images; pull is the default, so no callback URL is needed.

Network requirement: allow the runner outbound to the admin (HTTPS). No inbound rule to the runner is ever required in pull mode.

What you see in the admin UI

  • A runner with ephemeral badge in the runner list while it is alive.
  • The row disappears as soon as the runner deregisters (or the sweeper reaps it three minutes after last heartbeat).
  • Recent task history for that runner stays in the test-run history (Test Plans, Performance Tests) — task results are not tied to the runner row.

Troubleshooting

The runner row is not vanishing after the container exits.

  • Check the coordinator logs for Ephemeral runner DELETE failed warnings. The most common cause is a coordinator restart between the runner’s last deregister attempt and the next sweeper tick.
  • Wait three minutes. The sweeper runs every 60 seconds and reaps stale ephemeral rows whose last heartbeat is older than the MOCKARTY_EPHEMERAL_SWEEP_STALE_AFTER threshold (default 3 minutes).
  • Confirm EPHEMERAL=true was actually applied. The runner’s startup log line includes ephemeral=true when the flag is set.

The runner exits before finishing its work.

  • Decrease IDLE_TIMEOUT only when you know the runner gets at least one task within that window. A common pattern is to set IDLE_TIMEOUT=300s (5 min) — long enough for the CI pipeline’s previous step to enqueue work, short enough that orphaned containers vanish.
  • Setting IDLE_TIMEOUT=0s (the default) disables the idle watcher entirely — the runner keeps running until SIGTERM or its parent process kills it.

Deregister succeeded but the row is still there for 3 minutes.

  • This is the safety-net path (the deregister error fell through to the sweeper). The row is harmless during the window — the runner is no longer in the in-memory registry, so no tasks get assigned to it. You can shorten the window with MOCKARTY_EPHEMERAL_SWEEP_STALE_AFTER if your dashboards depend on clean state.

A runner joined but never picks up UI or mobile jobs.

  • Run the built-in readiness check on that machine — it needs no admin connection and changes nothing:

    mockarty-runner doctor
    

    It prints a truthful report of what THIS host can serve: ✓ ready works now, ◐ auto provisions on first use (the lightweight browser engine auto-installs), and ✗ needs action names what’s missing (for example adb not found).

  • Capabilities turn on by themselves. With just a token and an admin URL the runner probes the host and advertises every capability it can actually serve — no capability flags to hand-maintain. In particular:

    • Mobile (Android) auto-enables when a device is attached over adb (only authorized devices count). Override with RUNNER_MOBILE=true to make an ephemeral runner wait for a device, or RUNNER_MOBILE=false to opt out on a phone-less CI box.
    • RUNNER_BROWSERS=true asks for real browsers: the runner scans the Playwright browsers already installed and offers each one alongside the always-available lightweight engine (which downloads on first use). Fine-grained control stays available via RUNNER_UI_ENGINES.