Docs Autonomous Coder

Autonomous Coder

The Autonomous Coder turns one plain-language goal into working, deployed code. You describe what to build and point at a git repository — Mockarty decomposes the goal into sub-tasks, dispatches them to a fleet of coder runners, reviews every result with its own acceptance gate, reruns red work until it is green, and (optionally) deploys the accepted code to a target you configured. A human stays in the loop exactly where it matters: clarifying questions, and the deploy approval.

Open Autonomous Missions (/ui/missions?view=coder). The unified cockpit shows coder missions together with their runner capacity, job queue, evidence, economics, knowledge and per-component model routing.

The module is gated by the A2A / Autonomous feature. Deploy approval and delivery configuration additionally require an owner or admin role in the namespace.

Waiting for a runner

Jobs wait in the queue when all compatible runners are busy. A runner with one
job slot takes one job at a time; another compatible runner can take waiting work
as soon as it has capacity. Jobs requiring a different engine, label or testing
toolchain do not block compatible jobs further down the queue.

For push runners, registration and a completed job wake delivery. Heartbeats
also recover waiting work after an admin restart or an exhausted dispatch window;
this recovery is eventual, not an instantaneous scheduling guarantee. A refusing
runner does not prevent another compatible runner from receiving work. An
ambiguous delivery is not treated as proof that execution never started.

Job slots limit jobs, not the coding engine’s internal sub-agents.

Check capacity

The Capacity tab shows runners and the current job queue. For scripts, the
small capacity response gives runnerCount, onlineRunners, slots, busy
and queued for the selected namespace. It can be up to two seconds old;
slots counts online runners. A missing response means capacity is unknown.

curl -H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
  "$MOCKARTY/api/v1/coder/capacity?namespace=$NAMESPACE"

When you need recent job details, use /api/v1/coder/jobs?limit=50 (1–100).
Omitting limit preserves the complete existing list response.

Starting a mission

Click New mission and fill in:

  • Mission goal — plain language, e.g. “Add a REST endpoint /health that returns build info, with tests”.
  • Target git repository URL — the repo the coder clones and works in.
  • Engine (optional) — the coding engine the runner uses (opencode is the default).

Over MCP the same is one call:

{
  "tool": "coder_mission_start",
  "arguments": {
    "goal": "Add a REST endpoint /health that returns build info, with tests",
    "repoUrl": "https://git.example.com/team/app.git",
    "issueKey": "APP-42"
  }
}

issueKey ties the mission to a tracker issue: commit messages start with the key (for example, APP-42: ...), the issue panel links the mission, and the efficiency dashboard compares the issue’s estimate with the mission’s actual runtime. Mockarty does not add an auto-close verb to every sub-task commit; your merge policy decides when the issue closes.

The mission lifecycle

A mission moves through five phases, shown as a flow strip in the mission control panel:

Analyze → Plan → Build → Accept → Deploy

  • Analyze (optional, analyze: true) — the analyst produces requirements, an architecture decision record and a design document, publishes them to the wiki, files tracker issues, and authors test cases from the functional requirements. If the goal is ambiguous, the mission pauses with clarifying questions.
  • Plan — the goal (or the design) is decomposed into ordered sub-tasks.
  • Build — sub-tasks are dispatched to coder runners; each returns a branch and a structured report.
  • Accept — Mockarty starts a separate testing mission and executes the requested Mockarty functional, browser, load, fuzz, security, chaos or contract engines. A green receipt is accepted only for the exact namespace, mission, job, attempt, branch and commit, and only when every requested engine returned its own durable evidence artifact. The coder’s report and the optional model reviewer remain advice. A red independent verdict sends the sub-task back with concrete facts; this loop is bounded by the per-task attempt cap. If the independent route is unavailable, the result is kept as unverified and cannot authorize an automatic deploy.
  • Merge request — once every sub-task is accepted, Mockarty opens the merge (or pull) request for the branch it pushed and records its number, target, exact head and mergeability. A conflict or a changed head blocks the merge. The link and state appear on the mission. With an auto merge guardrail, Mockarty may merge only a conflict-free request whose head still equals the independently accepted commit and whose mission has no unverified jobs; the default remains human approval. Set openMr: "off" when your own automation owns requests. Re-running a mission never duplicates the same request, while a repaired commit receives a new exact effect identity.
  • Deploy — when a deploy target is set, the accepted code is delivered through SSH (copy artifacts to a host and run them), Compose (upload a docker-compose stack and docker compose up), or a CI pipeline (watch the company’s pipeline on GitLab, GitHub Actions, or Jenkins, and gate success on the run for the exact accepted commit). The immutable artifact is visible as mission.acceptedCommit; Mockarty refuses a mutable branch that no longer points at it. A GitOps target can commit manifests to a repository watched by Argo/Flux, but the current target deliberately remains unverified instead of claiming that the controller reconciled a healthy workload, and rollback remains unavailable until an exact GitOps deploy commit is recorded. A Kubernetes target applies the manifest from the workspace with server-side apply and gates success on kubectl rollout status for the named workload, so its verification is the rollout itself. Use CI with a health URL, SSH/Compose, or Kubernetes when the mission must finish with a verified deployment. Production-tier targets always pause for a human approval.

The execution safeguards apply to the whole mission, including recovery after an admin-node restart. maxAttempts defaults to 5 and accepts values from 1 through 20 per sub-task; invalid negative or larger values are rejected. A mission also has a durable active-time deadline and a mission-wide dispatch ceiling that failover cannot reset. Operator pause time does not consume the active deadline. Explicit human retries receive a fresh active-time window but are themselves bounded; automatic orphan adoption only resumes the remaining durable budget.

The mission control panel

Click any mission card (or the expand button) to open its control panel:

  • the phase flow strip with the current phase highlighted;
  • the plan with each sub-task’s live status and attempt count;
  • the timeline — everything the mission did, when, and why;
  • the deploy result with the endpoint link once deployed;
  • the mission’s exact metered consumption: model calls/tokens/cache, runner and tool resources, unpriced events, and the customer amount per currency.

The mission’s Economics tab keeps actual immutable-ledger consumption separate from the caps selected for the run. The page-wide economics summary reports spend, prefix-cache use and budget utilisation only when its accounting storage and roll-up queries are available. During an accounting outage the operational mission counters remain visible, while economics is explicitly marked unavailable; missing measurements are never displayed as zero spend.

Controls available from the panel, depending on state:

Control What it does
Pause / Resume Holds the mission at its next step; in-flight work finishes.
Cancel Uses the common durable Mission cancellation receipt, then stops the exact active Coder execution and queued sub-tasks (confirmation required). A 202 means the command is still reaching its owner or a remote child; it is not a false claim that all work has stopped.
Skip Force-accepts one sub-task. A reason is required and recorded in the timeline and the audit log.
Add sub-tasks Appends follow-up instructions to a live mission; each chains onto the latest pushed branch.
Retry Re-drives a failed, canceled or interrupted mission from its saved plan.
Restart from here Re-drives from a specific sub-task; earlier accepted deliverables are kept.
Reconcile deployment outcome Records whether an unverified deployment was actually applied. Only not_applied unblocks Retry.

When a retry is refused

If a mission stopped while its deployment outcome could not be verified — the
commit may have been pushed, the pipeline may already be running — Retry and
Restart from here are refused with a message naming the target. Re-driving
would apply that deployment a second time, and Mockarty will not guess what
happened on the other side.

Check the target yourself, then record exactly what happened. If the effect did
not land and retry is safe:

curl -X POST "$MOCKARTY/api/v1/coder/missions/$MISSION_ID/deploy-outcome" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"outcome":"not_applied"}'

Use applied instead when inspection proves that the original deployment did
land. In that case Mockarty settles a failed or interrupted mission as completed
and does not redispatch the deployment. An operator-canceled mission stays
canceled, while still recording that its deployment was applied. The outcome is
mandatory, recorded in the audit log with your name, and an exact repeated
statement is harmless; a contradictory second statement is refused. Only
owners/admins with deployment permission may make this decision. An agent uses
coder_mission_reconcile_deploy with the same explicit outcome. The CLI
equivalent is mockarty coder-delivery missions reconcile-deploy MISSION_ID not_applied.

For a CI pipeline target you do not have to inspect anything yourself: send
{"outcome":"provider"} and Mockarty asks the CI system (GitLab, GitHub or
Jenkins) about the exact pipeline, run or build it recorded when it dispatched
the deployment. A finished successful execution settles the mission as applied
from the provider’s own record; a cancelled or skipped execution settles it as
not_applied. A job that is still running is reported back so you can wait, and
a failed job is deliberately left to you — a failed deploy job may have applied
part of its steps. SSH, Compose, GitOps and Kubernetes targets have no such
record and keep the explicit applied / not_applied statement.

In a multi-node installation a control action may return the note “the command was queued and will be applied by its owner” — the mission is driven by another node and applies your command on its next tick (typically under a second).

The human in the loop

Missions that need you surface at the top of the page as attention cards, count into the Waiting for you indicator, and arrive as notifications in your inbox (and any channel you subscribed, e.g. Slack or Telegram):

  • Needs input — the analyst asked clarifying questions. Answer them in the card (or via coder_mission_answer) and the mission resumes.
  • Awaiting approval — a deploy is ready. Approve or deny from the card (or via coder_mission_approve). Only owners/admins can approve.
  • Paused — an operator deliberately held the mission. Resume it when the dependency is ready, or cancel it if the work is no longer needed.
  • Interrupted — the previous execution cannot safely continue after a lost owner or an ambiguous effect. Retry the active node to create a fresh fenced attempt, or cancel the mission. Mockarty never labels this state as failed or silently replays the old effect.

If the attention queue cannot be read, the page shows a retryable error instead of an empty healthy queue.

You also get a notification when a mission completes or fails.

Honest verdicts

If the independent testing mission could not run or could not return an exact typed receipt, the build is kept — but it is counted as unverified, shown as a ⚠ badge on the mission and in the Unverified accepts indicator. An automatic deploy is downgraded to requiring human approval in that case. Not checked is never presented as passed.

Who executes a sub-task

By default every sub-task is built by the coder-runner fleet. But one mission plan can weave in other capabilities — set a sub-task’s executor:

  • runner (default) — a coder runner writes the code.
  • agent — an internal Mockarty specialist drives the sub-task instead of writing code: a browser UI test (agentName: web_ui_tester), test-case authoring (test_planner), and so on. Its result is accepted (or reruns) exactly like a coder job.
  • remote — an external agent attached over A2A (Admin → Remote Agents), selected by the sub-task’s skillId and optional selector labels (e.g. region=eu). This is how you extend the fleet with third-party agent networks.

An unknown executor safely falls back to the runner. Over MCP, coder_mission_add accepts executor / agentName / skillId / verify / requiredChecks per task; the decomposer can also route a step to a specialist on its own.

How a sub-task must be proven

A sub-task can name the testing engine that has to confirm it, instead of asking for it in prose inside the prompt. Add verify to the sub-task:

{
  "prompt": "Add the /orders endpoint …",
  "verify": [
    {"engine": "functional", "ref": "orders smoke", "target": "http://localhost:8080", "note": "every request answers 2xx"},
    {"engine": "load", "note": "p95 under 300 ms"}
  ]
}

engine is one of functional, load, fuzz, chaos, contract, ui_test, test_case, temporal_probe, bot_scenario. An engine this installation cannot run is refused when the sub-task is dispatched and the reason appears in the mission timeline — it is never quietly promised. The coder receives the directive as concrete Mockarty tool calls it must quote in its report; where its toolset cannot reach the engine, it is required to say so rather than claim the check.

The coder reaches those engines through Mockarty’s own tools: the runner is handed the team-knowledge, mocking, API-testing, test-plan and deployed-system-observability tool groups, filtered by what your plan includes. An administrator can change that set with the MOCKARTY_CODER_MCP_GROUPS environment variable on the admin node (a comma-separated list of group ids, e.g. processing,mocker,tester,testplan,observability,perf); the group ids are the ones GET /api/v1/mcp/groups returns.

For deterministic local verification, add requiredChecks to the sub-task. Each check is an executable plus arguments, not a shell string:

{"requiredChecks":[
  {"name":"Go unit tests","args":["go","test","./..."],"timeoutSeconds":900},
  {"name":"Go vet","args":["go","vet","./..."]}
]}

The admin dispatches the task only to a runner that advertised the required toolchain. The runner executes the checks against the committed head before push. A non-zero exit blocks publication and returns a typed result (name, exit code, timestamps, and output digest); diagnostic text is bounded and redacted. Shell wrappers are refused because the queue cannot prove which toolchain they hide.

The same payload can be appended to a live mission with the CLI:

mockarty-cli coder-delivery missions add MISSION_ID --file follow-up.json

Go SDK

mission, err := client.CoderDelivery().AddToMission(ctx, missionID, mockarty.CoderMissionAddRequest{
    Tasks: []mockarty.CoderSubTask{{Prompt: "Run and fix the unit suite", RequiredChecks: []mockarty.CoderRequiredCheck{{Name: "Go unit tests", Args: []string{"go", "test", "./..."}}}}},
})

Python SDK

mission = client.coder_delivery.add_to_mission(mission_id, {
    "tasks": [{"prompt": "Run and fix the unit suite", "requiredChecks": [{"name": "Go unit tests", "args": ["go", "test", "./..."]}]}]
})

Java SDK

Map<String, Object> mission = client.coderDelivery().addToMission(missionId, Map.of(
    "tasks", List.of(Map.of("prompt", "Run and fix the unit suite", "requiredChecks", List.of(
        Map.of("name", "Go unit tests", "args", List.of("go", "test", "./...")))))));

Watching what you deployed

A deploy that answered “OK” is not a system that works. If your installation has an observability source configured, the coder — and you, and any agent — can ask the deployed system directly:

curl -s -X POST http://localhost:5770/api/v1/observability/query \
  -H "X-API-Key: mk_your_token_here" -H "Content-Type: application/json" \
  -d '{"source":"prometheus",
       "expression":"sum(rate(http_requests_total{code=~\"5..\"}[1m]))",
       "correlation":{"environment":"staging"}}'

The SDKs use the client’s namespace automatically. For a direct HTTP request,
add ?namespace=your-space to query a particular space; your account or API key
must have access to it. A namespace-bound key cannot query another space.

sources, err := client.CoderDelivery().ObservabilitySources(ctx)
evidence, err := client.CoderDelivery().QueryObservability(ctx, mockarty.CoderObservabilityQuery{
    Source: "prometheus", Expression: "up",
    Correlation: mockarty.CoderObservabilityCorrelation{MissionID: mission.ID},
})
sources = client.coder_delivery.observability_sources()
evidence = client.coder_delivery.query_observability({
    "source": "prometheus", "expression": "up",
    "correlation": {"missionId": mission["id"]},
})
Map<String, Object> sources = client.coderDelivery().observabilitySources();
Map<String, Object> evidence = client.coderDelivery().queryObservability(Map.of(
    "source", "prometheus", "expression", "up",
    "correlation", Map.of("missionId", mission.get("id"))));

For the CLI, mockarty coder-delivery observe sources lists the bound sources. Put the same JSON body as the REST example into a local file and run mockarty coder-delivery observe query --file query.json.

GET /api/v1/observability/sources?namespace=sandbox lists what that namespace can be asked, with the defaults it applies (the last 15 minutes, 200 rows). Over MCP the same pair is observability_sources and observability_query. Your account must be allowed to read the reliability module in the actual namespace being queried. If a POST body names a different namespace from the URL, permission is checked for the body namespace too; access in another workspace does not authorize it.

Every answer is bounded and read-only, and it comes back with a evidenceDigest — a sha256 over the query and its result — so a reading can be cited later instead of retold. At least one correlation field (mission, deployment run, release digest, test run, trace, environment) is required: evidence nobody can attribute is refused rather than stored. Correlation labels must be plain identifiers — a label shaped like a credential or personal data is refused with a validation error, never stored. Log lines that quote a credential come back with the credential masked in place; one noisy log line never voids the whole reading.

Nothing configured is a valid answer. An installation with no observability stack returns an empty source list and answers a query with source_not_configured. A configured source that fails returns source_error (502); repair it and retry. Neither response is evidence of health.

An administrator configures a source on the admin node:

Variable Meaning
MOCKARTY_OBSERVABILITY_PROMETHEUS_URL Prometheus HTTP API root, e.g. http://prometheus:9090. Unset = no source.
MOCKARTY_OBSERVABILITY_PROMETHEUS_TOKEN Bearer token, when the source requires one.
MOCKARTY_OBSERVABILITY_PROMETHEUS_ID Name this source is recorded under in evidence (default prometheus).
MOCKARTY_OBSERVABILITY_PROMETHEUS_NAMESPACE Exact namespace allowed to use this source (default sandbox). A source is never shared implicitly between namespaces.
MOCKARTY_OBSERVABILITY_LOKI_URL Loki HTTP API root. Enables bounded log reads.
MOCKARTY_OBSERVABILITY_LOKI_TOKEN / _ID / _NAMESPACE Optional bearer token, evidence connection name (default loki), and exact namespace binding.
MOCKARTY_OBSERVABILITY_TEMPO_URL Grafana Tempo HTTP API root. Enables bounded OpenTelemetry trace search.
MOCKARTY_OBSERVABILITY_TEMPO_TOKEN / _ID / _NAMESPACE Optional bearer token, evidence connection name (default tempo), and exact namespace binding.
MOCKARTY_OBSERVABILITY_SENTRY_URL Sentry HTTP API root. Enables bounded error/debug event reads.
MOCKARTY_OBSERVABILITY_SENTRY_TOKEN / _ID / _NAMESPACE Optional bearer token, evidence connection name (default sentry), and exact namespace binding.
MOCKARTY_OBSERVABILITY_SENTRY_ORGANIZATION / _PROJECT Required Sentry organization and project identifiers.

The namespace binding is a security boundary, not a label. Mockarty does not rewrite arbitrary PromQL to add a tenant matcher, so a configured Prometheus endpoint must contain only data that members of the bound namespace are allowed to read. Use separate admin-node installations for namespaces with different monitoring access.

After a coder deployment, Mockarty runs up to three behavioral health probes across a bounded window. On a failed probe it briefly reads the compatible observability sources for the mission namespace before rollback, while the failing candidate is still present. If the probes hold, Mockarty reads those sources before acceptance. Explicit error/fatal events or an unhealthy up metric also trigger rollback. A bounded repair/redeploy attempt is allowed only after the restored target passes its own health window; an unavailable rollback is reported as impossible, and an unverified rollback requires reconciliation before retry. The mission keeps source counts and frozen evidence digests, including a timeline record of diagnostic digests from a failed candidate before a successful repair replaces its result. Raw observability responses and credentials are not copied into repair instructions. No configured source or an empty/truncated answer remains not_evaluated; it is never rewritten as healthy.

Team skills

Beyond the coder’s built-in skills, a mission can pull in your team’s installed agent-marketplace skills so the coder works the way your team does. Pass skills (a list of skill names) to coder_mission_start — each named skill you have installed is materialised into the coder’s sandbox alongside its built-ins, and the coding engine loads it on demand. Only skills your namespace has installed and is licensed for are provisioned; a name you don’t have is quietly ignored. This is how you teach the autonomous coder your conventions (a house code-style skill, a domain-specific checklist) without changing the product.

Delivery config

Delivery config (page header) is where you describe how your company ships: the repositories the coder may use, the deploy targets (SSH / Compose / GitOps / CI pipeline / Kubernetes), protected branches, and infrastructure notes the coder reads as context. Credentials are referenced by secret-store id — they are never entered or displayed here. Target commands are screened against a dangerous-command list before the config is saved.

Only an owner or administrator can change this configuration, start a mission that may deliver, or approve its deployment. Add each allowed repository explicitly; the mission must use one of those exact URLs. Do not put credentials in a repository URL. Select an Approval policy for each target in the UI. A review written by an AI model alone does not authorize automatic deployment: without a separate, trusted quality check (AQC), Mockarty asks a person to approve the deployment.

GitOps currently records the manifest-repository effect but does not have an Argo CD or Flux status connection. A commit alone is not evidence that the cluster reconciled it, so the Delivery Config API rejects a GitOps deploy target before saving it. Keep GitOps repository metadata for planning if useful, but use a supported target with real verification for autonomous deployment until controller observation is available.

SSH and Compose targets must define healthCmd; CI targets must define healthUrl; a Kubernetes target is verified by the rollout of the workload it names and needs neither. The API rejects an unverifiable target before storing the configuration. The deploy-time check remains as protection for older stored rows. This is intentional: a completed upload, process start, or green build is not by itself proof that the delivered application is reachable and healthy.

For GitLab CI, GitHub Actions and Jenkins, Mockarty retains the provider’s exact, credential-free execution identity as soon as the provider accepts the launch. A cancellation can therefore find the same pipeline, workflow run, queue item or build after an admin-node restart. An accepted cancel request is still not treated as completion: Mockarty waits for that exact provider execution to become terminal. If correlation, lookup or terminal confirmation is ambiguous, the mission stays unknown and requires reconciliation; Mockarty does not silently launch the deployment again.

The default Compose start (docker compose up -d) has an exact bounded inverse against the same target directory and Compose file: if verification is cancelled or fails, Mockarty runs docker compose down even though the mission context is already cancelled. If you provide a custom startCmd, also provide its exact rollbackCmd; Mockarty does not guess how to undo an operator-defined command.
After the default Compose rollback, Mockarty checks three times that the same stack has no containers. If it cannot confirm this state, the rollback remains unverified and automatic retry is held. A custom rollback is checked with healthCmd; configure it to check the intended restored state.

The Product autonomy settings section sets two defaults for every
autonomous mission in the workspace: the guard policy (pause the mission
and ask a person, or warn and continue) and the run window in minutes (a
mission running past it hits the time wall). A single product can override
them: an agent writes that product’s own delivery-config group via
coder_delivery_config_set with productId (or over the REST API), and it
wins over the workspace default for that product’s missions. The layer order:
a mission-level value wins, then the product’s, then the workspace’s, then
the instance’s.
The product context chips on the missions page — or the
mission_settings_effective agent tool — show each selected setting together
with the level that set it and runtimeApplied. Journal retention is applied as
a live namespace/instance cleanup policy, not as a product or per-mission
override. Configure it in Autonomous Missions → Limits and settings, leave a field blank
to inherit, or use mockarty-cli autonomy settings set. Active legal holds
preserve matching coder missions, jobs, intake evidence and journal rows even
after the normal horizon expires.

The CI/VCS section — {"system":"gitlab","baseUrl":"https://gitlab.company.internal","credRef":"…"} — is what lets Mockarty talk to your forge: it opens the merge request for a finished mission and, with autoReviewMRs: true, reviews every request your team opens. Add "openMr":"off" to keep request creation to yourself.

Each target carries a small JSON config. Examples:

  • SSH / Compose — {"host":"deploy.example.com","user":"deploy","targetDir":"/opt/app","mode":"compose","composeFile":"docker-compose.yml","hostKey":"…","healthCmd":"curl -sf localhost:8080/health"}
  • CI pipeline — {"system":"gitlab","baseUrl":"https://gitlab.company.internal","project":"team/app","healthUrl":"https://app.example.com/health"} (system is gitlab | github | jenkins; the branch defaults to the delivered one). GitLab uses the exact pipeline ID returned by the trigger. GitHub requires workflow, a workflow input named mockarty_delivery_id, and run-name: ${{ inputs.mockarty_delivery_id }}; Mockarty correlates that run name and the accepted commit. Jenkins follows the queue item returned by the trigger to its exact executable number and checks MOCKARTY_COMMIT. Mockarty waits for that bound run and the health check before the deploy counts as done. Put the CI API token in the target’s credential reference; inline tokens in JSON are rejected.
  • Kubernetes — {"context":"prod","namespace":"payments","manifestPath":"deploy/app.yaml","workloadKind":"deployment","workloadName":"payments-api","rollbackRevision":7,"timeoutSeconds":300}. workloadKind is deployment | statefulset | daemonset; manifestPath is relative to the delivered workspace and at most 4 MiB; rollbackRevision is required and must name an existing rollout revision, because that is what a rollback returns to. fieldManager defaults to mockarty, timeoutSeconds to 300 (10-900). The kubeconfig goes in the target’s credential reference; set "inCluster": true instead when Mockarty runs inside the cluster it deploys to. Rollback runs rollout undo to that exact revision and then re-checks the rollout — an unconfirmed rollout is reported as unknown, never as a successful rollback. A Kubernetes target needs kubectl on the machine running Mockarty; without it the deploy is reported as not having reached the cluster, not as an uncertain outcome.

For SSH and Compose, targetDir must be absolute and must not resolve to the filesystem root (for example, /opt/app/../.. is refused).

The REST surface is /api/v1/coder/delivery-config and /api/v1/coder/missions. The Go, Python, and Java SDKs expose it as CoderDelivery() / coder_delivery / coderDelivery(). The CLI mirrors the bounded operations:

mockarty-cli coder-delivery config get
mockarty-cli coder-delivery config put --file delivery-config.json
mockarty-cli coder-delivery config delete --product-id PRODUCT_ID
mockarty-cli coder-delivery missions start --goal "Ship the accepted commit" --repo https://git.example/app.git --target staging --product-id PRODUCT_ID
mockarty-cli coder-delivery missions list
mockarty-cli coder-delivery missions get MISSION_ID
mockarty-cli coder-delivery missions add MISSION_ID --file follow-up.json
mockarty-cli coder-delivery missions approve MISSION_ID
mockarty-cli coder-delivery missions deny MISSION_ID

The same start request is available through REST and every supported SDK. The
repository URL must exactly match a repository in the selected Delivery Config;
productId selects a product-specific config when one exists.

REST

curl -X POST "$MOCKARTY_URL/api/v1/coder/missions?namespace=default" \
  -H "X-API-Key: $MOCKARTY_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"goal":"Ship the accepted commit","repoUrl":"https://git.example/app.git","deployTarget":"staging","productId":"PRODUCT_ID"}'

Go SDK

mission, err := client.CoderDelivery().StartMission(ctx, mockarty.CoderMissionStartRequest{
    Goal: "Ship the accepted commit", RepoURL: "https://git.example/app.git",
    DeployTarget: "staging", ProductID: "PRODUCT_ID",
})

Python SDK

mission = client.coder_delivery.start_mission({
    "goal": "Ship the accepted commit",
    "repoUrl": "https://git.example/app.git",
    "deployTarget": "staging",
    "productId": "PRODUCT_ID",
})

Java SDK

Map<String, Object> mission = client.coderDelivery().startMission(Map.of(
    "goal", "Ship the accepted commit",
    "repoUrl", "https://git.example/app.git",
    "deployTarget", "staging",
    "productId", "PRODUCT_ID"
));

Models and runner capacity

Open Capacity in Autonomous Missions. The same panel shows connected runners, available slots, queued jobs and the model assigned to each component — decomposition, acceptance, code review, and the build model runners use. A workspace owner can add its write-only model profile directly there and assign it immediately; installation-wide profiles remain administrator-managed. All traffic runs through Mockarty’s LLM gateway: it is masked by the guardrails, metered into the usage ledger, and switchable centrally — change the assignment and every caller moves to the selected model.

Runner output and recovery

Each execution uses a separate private working directory. A retry does not overwrite files retained from an earlier attempt. The runner keeps unpublished commits and failed work after checkout for recovery. Successful published work, or a successful run with no changes, can be cleaned up after its terminal report has been saved outside that directory.

For one-shot CLI runs, use --report with a path outside the working directory. Without an external report destination, the working directory is retained. --keep-workspace also disables automatic removal. Retained directories consume disk space and may contain sensitive task files or logs: restrict access and recover any needed code and logs before removing them. They are not an automatic remote artifact backup.

Undelivered results also consume runner capacity. If durable result storage fails, the runner stops admitting new work rather than discarding results to make room. Restore storage and connectivity before retrying work; do not delete pending result files as a disk-cleanup shortcut.

Measuring the payoff

Two dashboard data sources (Dashboards → add widget → Autonomous Coder) quantify the autonomous work:

  • Coder mission LLM cost — tokens per mission over a period, top consumers first.
  • Coder missions: human-hours saved — the linked issues’ estimates minus the missions’ runtime to completion. The period is based on when a mission completed, so later updates do not move it to another period. Missions without a linked estimate contribute their actual runtime but no assumed human estimate; the net value can be negative.

Prometheus counters (/metrics) cover missions by status, durations, reruns and unverified accepts for your own monitoring stack.

MCP tools

The full mission lifecycle is drivable by agents: coder_mission_start, coder_mission_status, coder_mission_answer, coder_mission_approve, coder_mission_pause / resume / skip / add / retry / restart_from / reconcile_deploy / cancel, coder_mission_diagnose, plus coder_analyze, coder_review_mr, coder_ingest_artifact and the delivery-config pair.