Docs Autonomous Missions

Autonomous Missions

Autonomous Missions is the single place where all of Mockarty’s autonomous work is visible and controlled: what you started by hand, what an agent submitted, what arrived from your message queue, and what a schedule fired at night. One page answers three questions: what is running, what is waiting for a person, and what it costs.

About URLs in the examples: the examples use localhost:5770 as the Mockarty address. If your instance runs elsewhere, replace it with the real address.

The section is available on plans that include the AI-agent capabilities. If the menu entry is not visible, the feature is not part of your current plan.


What a mission is

A mission is one autonomous task described as a goal in plain words. You do not need to choose a specialist executor. Mockarty makes a plan using only the capabilities available in your workspace, such as analysis, coding, testing, security and delivery. You can inspect its steps and progress on the mission card. The original goal and each plan revision stay in the mission history.

Every mission always shows:

  • state — queued, running, paused, waiting for data, waiting for approval, done, failed, cancelled;
  • origin — the UI, an agent, a schedule, a message queue, a task from another system;
  • product and service — whose work it is;
  • spend — tokens used, and how much of the budget is left.

Attach original documents and designs

In the new mission wizard, select a product in Context, then choose files under Documents and design materials. You can attach up to 16 nonempty files: images and PDFs at most 16 MiB each, all files at most 64 MiB combined, and text documents at most 64 KiB combined. Uploads complete before the mission starts. If starting fails, retrying with the same product and selected files reuses successful uploads in the open wizard.

Image analysis requires the selected LLM profile to support vision. PDF analysis also requires a provider connection that supports PDF input. Providing originals to OpenCode does not by itself verify that its selected model can interpret them; visual fidelity must be checked through independent product acceptance before treating the result as verified.

To upload from a script, send one multipart file to POST /api/v1/missions/materials?namespace=default&productId=PRODUCT_ID. The 201 response includes material metadata and a reference with kind, id, digest, and revision. Include that reference in the artifacts array of POST /api/v1/missions, using the same namespace and product. Keep the returned digest and revision unchanged.

curl --fail-with-body -H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
  -F 'file=@design.png;type=image/png' \
  'http://localhost:5770/api/v1/missions/materials?namespace=default&productId=PRODUCT_ID'
mockarty-cli autonomy missions upload-material design.png --namespace default --product-id PRODUCT_ID

Go

result, err := client.CoderDelivery().UploadMissionMaterial(ctx, "default", productID, "design.png", "image/png", file)
// Use result.Reference in the mission's artifacts array.

Python

result = client.coder_delivery.upload_mission_material(product_id, "design.png", content, "image/png")
# Use result["reference"] in the mission's artifacts array.

Java

var result = client.coderDelivery().uploadMissionMaterial("default", productId, "design.png", content, "image/png");
// Use result.get("reference") in the mission's artifacts array.

Where autonomous coding can work

Coding tasks run in the workspace selected for the mission. They cannot read or
write outside that workspace, and the coding tool cannot approve its own
requests for extra access. Attach the needed files before starting. If the
mission needs a product choice or clarification, answer it in the mission
conversation.

Mission history

Open a mission to see its steps, tool results, questions, your answers, and
outcome in one timeline. Your reply appears beside the question under your
name.


The product is the page’s context

The product is chosen in the section header, and everything else follows from it: which repositories and credentials the coder uses, where delivery goes, which schedules are shown, which services the filter offers.

Inside a product, a task can belong to a service — one repository or component out of a hundred. It is a separate field rather than a separate product: a hundred repositories of one team stay one product, otherwise the product list turns into a junk drawer.

Triage: missions with no product

When a task arrives from outside, Mockarty works out whose it is: first from what delivered it (token, connector), then from the product named in the message, then from the product’s intake rules, then from the repository or host address. If nothing matches, the mission lands in Triage and waits for a person. If two products match equally, Mockarty does not guess — it names both and asks.

Such a mission’s card has a Bind button: pick the product (and the service, if relevant). The “and remember the rule” checkbox writes this task’s identifying detail into the product’s intake rules, so the next messages like it bind themselves. When the task carries no recognisable detail, no rule is stored: a rule that matches everything would pull in other products’ work.

The product must already exist in the current workspace. UI, REST, SDK and MCP requests all verify that boundary before a mission or an intake rule is changed; an unknown or foreign-workspace identifier returns 404.

The behaviour on ambiguity is configured per workspace:

Mode What it does
ask (default) Accepts the work, asks about the product — the mission waits in Triage.
create Creates a product from the name in the message.
strict Rejects a task that cannot be attributed to a product.

Starting a mission

  1. Click Start a mission and describe the goal in plain words.
  2. Choose the product and, if useful, a service. Add context or documents the work needs.
  3. Set spending and time limits if the defaults are unsuitable. The run window
    defaults to eight hours; when it expires, the mission pauses for a decision.
  4. Review the goal, context, permissions, and limits, then press Start.

Mockarty uses only the tools available to that mission. If the work cannot be
planned with those tools, you see an error before execution starts. If a tool
is temporarily unavailable, the mission card shows that delivery is delayed
and when another attempt is expected.

If the product or workspace settings change while you review the mission,
Mockarty asks you to review the updated settings before starting it.

Pause, cancel, and retry

Use the mission card to pause, resume, or cancel a mission. You can also retry or skip its active step. Add a reason when cancelling if other people need to understand the decision. A paused mission can resume; Retry starts a fresh attempt for an interrupted step.

After cancellation, check the result before assuming external work has stopped. Mockarty may still be waiting for a runner, broker worker, or deployment system to acknowledge the command. The response distinguishes acknowledged, refused, pending, and unknown work. If executionBindingsAvailable=true, inspect each row of executionBindings[] for the exact external execution and its state. An empty list means no external child was bound. If availability is false, inspect the executor directly. A late step result cannot overwrite a cancellation decision.

For API and MCP automation

API and MCP clients can attach an Idempotency-Key (or idempotency_key in the MCP tool). Reuse that key after a lost response: Mockarty returns the original control receipt and does not create a second operator command. A cancellation or skip reason is returned as control.reason on every exact replay. Mission answers are deliberately different: the answer stays in the mission record and reaches the executor, but is omitted from public control receipts, mission-control events and audit metadata so a possibly sensitive value is not copied into evidence surfaces. Delivery is retried automatically only when the executor can safely deduplicate the same intent. A 202 with control.outcome=pending means the command is still in bounded delivery or local recovery; retry the same key to read and advance that receipt safely. A definite refusal is returned as 409. If the connection failed after dispatch and the outcome cannot be proved, the response is 202 with control.outcome=ambiguous; Mockarty does not guess and does not automatically replay a potentially destructive command. If Mockarty still cannot finish the durable update after bounded recovery, the API returns 503 with control.outcome=interrupted; the mission stays blocked from conflicting updates and the receipt is queued for the audit log.

Cancel from automation with the same durable contract:

cURL

curl -X POST "$MOCKARTY_BASE_URL/api/v1/missions/$MISSION_ID/cancel" \
  -H "Authorization: Bearer $MOCKARTY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: incident-42-cancel" \
  -d '{"reason":"task is no longer relevant"}'

CLI

mockarty-cli autonomy missions cancel "$MISSION_ID" \
  --reason "task is no longer relevant" \
  --idempotency-key incident-42-cancel

Go

receipt, err := client.AutonomousMissions().Cancel(ctx, missionID,
    mockarty.MissionCancelRequest{
        Reason: "task is no longer relevant", IdempotencyKey: "incident-42-cancel",
    })

Python

receipt = client.autonomous_missions.cancel(
    mission_id,
    MissionCancelRequest(
        reason="task is no longer relevant",
        idempotency_key="incident-42-cancel",
    ),
)

Java

MissionControlResponse receipt = client.autonomousMissions().cancel(missionId,
    new MissionCancelRequest()
        .reason("task is no longer relevant")
        .idempotencyKey("incident-42-cancel"));

Answer a waiting mission with the same retry identity. The reply is omitted from control.reason by design:

cURL

curl -X POST "$MOCKARTY_BASE_URL/api/v1/missions/$MISSION_ID/answer" \
  -H "Authorization: Bearer $MOCKARTY_API_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: incident-42-answer" -d '{"answer":"use the sandbox account"}'

CLI

mockarty-cli autonomy missions answer "$MISSION_ID" --answer "use the sandbox account" \
  --idempotency-key incident-42-answer

Go

receipt, err := client.AutonomousMissions().Answer(ctx, missionID,
    mockarty.MissionAnswerRequest{Answer: "use the sandbox account", IdempotencyKey: "incident-42-answer"})

Python

receipt = client.autonomous_missions.answer(mission_id,
    MissionAnswerRequest(answer="use the sandbox account", idempotency_key="incident-42-answer"))

Java

MissionControlResponse receipt = client.autonomousMissions().answer(missionId,
    new MissionAnswerRequest().answer("use the sandbox account").idempotencyKey("incident-42-answer"));

For an ambiguous or interrupted receipt, open the mission card and choose Reconcile receipt. First inspect the executor run itself, then record either Command was applied or Command was not applied and describe the evidence. Mockarty never sends the original command again during this recovery; it resumes only its local fenced update. The same operation is available as POST /api/v1/missions/{id}/controls/{controlID}/resolve with {"resolution":"applied|refused","reason":"what was checked"}, and as the MCP tool mission_control_resolve. It requires deploy permission. An exact retry is idempotent; a different decision conflicts. mission_get returns the blocking public control receipt after a reload. Public receipts use stable error codes and do not expose provider, transport, or database diagnostics; administrators can investigate the protected audit trail and server logs.

Portable mission archives

An at-rest mission can be exported together with its immutable Brief and complete journal, then restored into another compatible Mockarty instance. The archive stays bound to its original mission ID and namespace; restore never remaps either authority. Queued and running missions cannot be exported. A missing immutable Brief, a journal gap, a changed digest, or any identity conflict is rejected without a partial restore. Replaying the exact same archive is safe and returns created=false. Large artifacts remain referenced by identity and are not copied into the archive.

mockarty-cli autonomy missions archive export "$MISSION_ID" --file mission.mockarty-archive.json
mockarty-cli autonomy missions archive restore --file mission.mockarty-archive.json

The equivalent SDK calls are ExportArchive / RestoreArchive in Go, export_archive / restore_archive in Python, and exportArchive / restoreArchive in Java. The file contains mission metadata and journal payloads; protect it as operational evidence.

Retries of tools with external effects

For durable agent work, Mockarty records an execution receipt before starting a tool. If a completed call is replayed after a restart, the saved result is returned without running the tool again. Tools that support deduplication receive the same idempotency key on an uncertain retry. Safe read operations may be repeated normally.

An arbitrary external service cannot be made exactly-once by Mockarty alone. If a tool has no deduplication contract and the server stopped after dispatch but before saving the result, the step result starts with DELIVERY WARNING. The action may already have happened and the recovery attempt may have applied it again. Inspect the target system before any manual retry. If the receipt itself could not be saved, DURABILITY WARNING means automatic repetition is unsafe until the target is reconciled. These warnings are part of the mission trace and remain visible after restart.


When a mission waits for you

Two states mean the work cannot continue without a person, and both are resolved right in the card.

Waiting for data. The mission asked a question — a test account, an access detail. The question is shown in the card; Answer records a durable command before handing your reply to the engine, and work resumes only after the engine accepts it. Repeating the same idempotency key returns the original receipt instead of sending a second answer. The answer stays on the mission record but is not copied into the public receipt or audit metadata.

Waiting for approval. The mission reached a step that needs a human decision (a deployment, for example). Approve and Reject are in the card; when rejecting, give a reason — it stays in the mission’s journal. Approval requires the delivery permission.

The Waiting for you counter in the header covers all such missions; clicking it filters the list.

What the mission did and what it left

The mission card answers not only “what is happening” but “why did it end that way”.

Execution trace — the steps the engine took: what it tried, how it ended, what it cost in tokens, and a link to the step’s source. A failed step is marked — that is where you start. When a step was a question, the human answer sits next to it: “asked X → answered Y” only reads as a pair.

What the mission left behind — the mocks, collections, test cases, reports, anomalies, the merge request it opened, the delivery result. That turns “done” from a claim into a checkable fact, and ties the run to the product’s entities.

The card also carries a “Deep dive in the engine” link that opens the execution flow on the page of the engine running the work: the step tree, the agent tasks, the gathered context. The section answers “what and why”; the engine’s page answers “how exactly” — there is no reason to duplicate it here.

Execution strategies

Ask first — the tester shows its plan and waits for you; the coder stops on clarifying questions and before a deploy. This is the default for missions started here: you are around and see what happens.

End-to-end — the mission does everything itself, with no stops. This is the default for work arriving from outside: a system usually stands behind it, not a person, and there is nobody to wait for.

Recon only — the mission looks and reports, changing nothing. The coder cannot run this way: it exists to change code, so a mission with a coder node set to recon refuses to start and says why — that is honest, unlike making changes nobody asked for.

Coder execution safeguards

The Coder page and the Autonomous missions page use the same start window. It keeps the product, repository, budgets, team playbooks and execution controls together, so a task does not become less safe when it is started from a different menu.

  • Maximum coder attempts is 1–20 and includes the first attempt. The default is 5. Failed work is retried with spacing between attempts and stops at the selected limit; cancellation and the mission deadline stop it before another attempt is sent.
  • Review depth can use the product setting, run the coder only (fast), or add specialist review (thorough). Unknown values are rejected instead of silently changing the requested review.
  • Delivery target is optional and must match a target configured for the selected product. Leave it empty to build and review without deployment. Protected targets still require their configured approval.

When the Processing module is available for the user and namespace, a coder mission can use the analyst/architect stage before implementation. Without Processing, the coder still decomposes and executes the task, but does not receive that separately licensed capability.

The team’s playbooks

The section ships with ready playbooks for four roles: the analyst (a requirement with checkable criteria; what a change would break), the architect (recording a decision; contract-first design), the developer (a change taken to proof; a merge-request review), the tester (a bug turned into a permanent regression case; release readiness on five points) — plus one for everybody: never report “done” without an observation that supports it.

Pick them right in the start-mission window: the ones you tick travel with the task, and the executor follows them instead of inventing its own order. Pick none and it works as usual. Your own playbooks are authored where the rest of the agent skills live, and one with the same name overrides the shipped one.

Which model does what

The Capacity tab answers that too: every mission node has its own model, and separate rows assign the model for the course check and for the coder’s internal steps (splitting the task, accepting a sub-task, reviewing a merge request, developing on the runner). Unassigned means the default model — so everything works out of the box with no configuration at all.

What the coder can build and check with

The coder works in a container, and that container ships a working baseline of the three most common stacks from the start — so it can not only edit files but build and verify what it did:

  • Go — the full toolchain: build, tests, go vet;
  • Python 3 — the interpreter, pip and venv;
  • Node — frontend work and npm tooling;
  • build tools — make and the C/C++ compilers (native Python modules and some Go dependencies do not build without them);
  • a real browser — the coder opens the page it produced, reads the console, clicks what it built and confirms the behaviour is the one it wrote;
  • everyday shell — git, curl, jq, ripgrep, unzip, rsync, an SSH client and kubectl.

The runner card’s Capacity section lists the tools it can use and shows a “no toolchain” warning when none are available. Check it before assigning coding work.

Your stack is not on the list? Add it as a layer on top — the base image deliberately leaves out the JDK, .NET, Rust and PHP: each adds hundreds of megabytes for a fraction of the work.

FROM mockarty/coder-runner:latest
USER root
RUN apt-get update \
    && apt-get install -y --no-install-recommends openjdk-21-jdk maven \
    && rm -rf /var/lib/apt/lists/*
USER coder

Build that image under your own name and point the runner deployment at it. The new tools show up in the advertisement by themselves — the runner probes its environment at start-up, with nothing extra to configure.

What the coder already knows about your stack

Before the first edit, the coder reads the repository itself — go.mod, package.json, pyproject.toml, pom.xml or build.gradle at the root and one level down — and picks the matching engineering playbook: Go, Node/TypeScript backends, Python (FastAPI, Django), Java/Spring, or a web UI framework (React/Next, Vue/Nuxt, Svelte, Angular). A monorepo with a Go API and a React frontend gets both. Each playbook is the review floor a senior engineer of that stack expects: error handling, validation at the boundary, layering, transactions, tests that fail without the change, and the stack’s own linters and test commands. When the repository has a UI, the coder also holds the frontend bar — keyboard and screen-reader access, contrast in both themes, every interactive state, responsive layout down to a phone width — and the copy rules for buttons, errors and empty states.

Your own conventions win: an AGENTS.md, CLAUDE.md or CONTRIBUTING.md in the repository and its configured linters override the playbook wherever they differ. Nothing is written into your repository — the playbooks live in the checkout the coder works in and are excluded from commits.

How thoroughly the coder works

A product chooses: fast — the coder alone, or thorough — specialists (a reviewer, an interface designer) join every sub-task and the task carries guidance on using them. Thorough pays off on risky changes and is wasted on one-line fixes, so the choice lives with the product rather than being switched on for the whole installation.

Leave it unset and the coder works as its runner is configured — everything works out of the box without this setting.

Is the mission still on course

The section shows by itself when something looks wrong with a run: the mission repeats the same step, failures come one after another, the budget is being spent with nothing produced, or the mission counts as running and has not done anything for a long time. These show at the top of the card — you need to see them before the budget runs out. They stop nothing: the decision stays yours.

Those signs see the shape of the work, not its meaning. A mission can march briskly from step to step and work on the wrong thing entirely. That is what Check course is for: the model reads the goal, the chain, what has been produced and the last steps, and answers whether the mission is doing what it was asked to. The answer — “On course” or “Course in doubt”, with concrete reasons — stays on the card, so the next person and any agent see it too.

The check stops nothing and costs one model call, so it runs on your command rather than by itself. It earns its keep on cheaper models: drifting off the task is their usual way of failing, and they do not report it.

A chain node can be retried or skipped — the command goes to the engine that runs the work. Only the active node can be skipped; a future node is never allowed to stop or fence the active attempt. When the engine cannot do it (the autonomous tester picks its own checks, for instance), the section says so plainly instead of pretending it skipped.

When the product does not answer

A mission that was given a product address first checks that something answers there, and only then starts working. The check is patient with a deployment that is still coming up: a few attempts with pauses, about half a minute in total.

No answer means the mission stops and asks instead of testing. The card shows a question with the address and the verbatim reason (“connection refused”, “connection timed out”), and no budget is spent: zero tokens, zero steps. That is deliberate — an unreachable product is not a verdict about its quality, it is a reason to supply a working address or fix the deployment.

You answer in the same card. If your answer contains an address — “redeployed, check https://stage.example” — it becomes the mission’s new target and the work continues from there. Otherwise the answer would stay words while the next step went back to the dead address.

The address comes from the task (the productUrl field or a product_url reference), or, when the mission is bound to a product, from that product’s addresses in the catalogue. A mission with no address — one about existing test cases, say — is not gated and does not wait: there is nothing for it to open.

Testing as a restricted account

Some checks only make sense as a particular user: “a read-only role must be refused clearly”, “this release may not change anything”. Give the mission that account instead of describing it in the goal.

  1. Put the account’s token in Secrets Storage.
  2. Create a connection whose endpoint is the product, whose target policy covers the product’s paths and methods, and whose allowedOperationIds include <adapter>.target_read and <adapter>.target_write (for example shop.target_read, shop.target_write).
  3. Submit the task with the product address and a restricted_connection_ref reference naming the exact revision:
{
  "goal": "The account is read-only. Check that the interface refuses creating items clearly.",
  "autonomy": "auto",
  "options": ["ui"],
  "productUrl": "https://shop.example.com",
  "contextRefs": [{"kind": "restricted_connection_ref", "value": "shop-reader@1"}]
}

What changes for such a mission:

  • It reaches the product only as that account. API calls go through the target_request tool, and a browser session is opened for the product and acts as the account — the credential is added by Mockarty and never reaches the model, the browser runner or the report.
  • Every request to the product leaves a receipt. A 403 is recorded as permission_denied and treated as a result to report, not as a broken tool. A write is a real attempt, so a mission told not to write can be checked afterwards: its receipts show whether it tried.
  • Tools that would reach the product another way — ambient HTTP requests, saved UI tests, scanners, load, delegation to other agents — are not available to it.
  • Checks that would otherwise need your approval as destructive (security, load) do not stop for it: without their engines the most they can do is write to the product as the account.

The mission is refused before any work, with the reason in the card, when the reference does not resolve to that exact revision, when the connection does not allow target_read, when the product address is not the connection’s own endpoint, or when the task names more than one restricted account. The browser part needs a web session on a real browser engine (Chromium, Firefox or WebKit); mobile sessions and saved sign-in states cannot be combined with a restricted account.


Tasks from other systems

A mission does not have to start in the UI. It can be submitted by an agent over MCP, by a neighbouring system as an A2A task, by a message on a queue (NATS, Kafka, RabbitMQ), by a tracker webhook, or by a schedule. Any such task shows up in the section immediately — with its goal, origin, planned graph and budget — and is then driven exactly like one started by hand.

The ordinary contract is goal-first: Mockarty plans the complete work from the goal and the authorized capability catalogue. When a task also lists the checks it wants (options, for example ["api", "security"]), every listed check the task may use becomes part of the plan — the planner can add more, but it cannot drop one. Listed checks never create a second mission system or bypass the common planning, evidence and cancellation authority.

When the goal is vague, the mission asks for what it is missing and waits — the question is visible in the card, and the answer is given right there.


Schedules

The Schedules tab holds a product’s recurring missions: a nightly check, a weekly load run, a regular security audit.

A schedule carries its own prompt: you write what exactly to do (“run card payment and refunds, find the regression after the release”), and that text becomes the mission’s goal. Besides the prompt you set a name, a cron expression (0 3 * * * — every night at three; shorthands such as @daily work too) and, optionally, the service. Mockarty builds the execution route from the goal when the schedule fires; you do not select a leading node.

A fired schedule becomes an ordinary mission with the “Schedule” origin — in the morning it is in the common list next to the rest, with its goal, planned graph and spend. While a schedule’s previous run is still open, the next one is not created: two concurrent runs of the same task only get in each other’s way.


Workflow builder

The Builder tab assembles a repeatable workflow as a versioned graph of steps from the capability catalogue, checks it with a dry run (capabilities, connections, secrets, cost upper bound) and publishes an immutable version. Publishing does not start a mission. Step by step: Versioned Workflow Definitions.


Triggers and Knowledge

The Triggers tab answers “what starts work here”: the product’s schedules (create one with a goal, a cron and a product — the run then happens by itself), the workspace’s inbound listeners, and an honest count of where missions actually came from in the last 7 days, taken from the records rather than the configuration. Configure connections opens the connector editor directly in the missions cockpit. It can create inbound NATS, Kafka, RabbitMQ, HTTP, Telegram, Slack and issue-trigger connectors, plus outbound webhooks; edit their settings, disable or re-enable them, and remove them. Type and direction cannot be changed after creation. Disabled connectors remain visible for re-enabling. The editor never displays a saved secret or destination URL: leave either field blank when editing to keep its existing value.

An inbound HTTP connector uses an HMAC secret and its copied /api/v1/intake/{id} address. Telegram uses a webhook secret token at /api/v1/public/intake/{id}/telegram; Slack uses a signing secret at /api/v1/public/intake/{id}/slack. The editor provides the matching copyable address for each. NATS, Kafka and RabbitMQ use their broker configuration and credentials; they do not verify this webhook secret. An issue trigger needs a status or assignee rule and runs from an authenticated issue webhook. For these non-webhook connectors, copy the connector ID rather than a webhook address. The connector API also accepts GET /api/v1/autotester/connectors?include_disabled=true for an authorized reader to manage disabled entries; without that flag, the list still returns only enabled entries. List and detail responses do not reveal stored secrets.

An outbound webhook needs an absolute HTTP(S) destination URL and a signing secret. It sends meaningful mission events to that address; choose Include every progress event only if the receiver needs step-by-step updates. The receiver can verify X-Intake-Signature against the JSON body. A destination on a private or local network may be refused by the server’s network policy.

The connector detail view shows only known non-secret configuration fields. Unknown values and broker addresses that may contain credentials are masked; leaving a masked field blank keeps its saved value, while entering a replacement updates it. The detail response includes updated_at; send it as expected_updated_at in a PUT to protect edits made in an older window. If another editor changes the connector first, saving returns 409; close and reopen the editor to load the latest settings. A connector deleted during editing returns 404 and is not recreated.

The Knowledge tab is what the product knows and what the installation has learned. Product facts (an example request, a domain invariant, a known issue, an environment note) are added by people and agents and feed the context the autonomous tester starts from; a stale fact can be deleted — future runs will no longer see it. Lessons are searched by meaning across the whole installation’s experience. The same surfaces are available to agents as qa_product_facts_list / qa_product_facts_add / qa_product_fact_delete, experience_search, and mission_triggers.

Deployment incidents

When something a mission deployed misbehaves, record it as an incident instead of a chat thread. Open the Incidents tab, choose the product in the header and press Open incident. Pick the mission; the deployment run is optional. The artifact, environment and product come from that deployment, so an incident always points at exactly what shipped. A mission that has not deployed anything cannot carry an incident.

Each incident collects hypotheses — possible causes, each with a short code (db_timeout, config_drift) and the reason someone suspects it. People and agents both propose hypotheses.

Attach evidence backs a hypothesis and marks it as supporting, refuting or context. There are two ways:

  • Reference to a material — where the evidence is: run://<test run id>, report://…, log://…, metric://…, trace://…, event://…, artifact://… or deployment://…. This always works, including without any observability stack.
  • Observability query — when observability sources are configured for the namespace (see “Watching what you deployed” in Autonomous Coder), pick a source, what to read and a window (15 minutes by default). The query is tied to the incident’s mission, deployment and artifact automatically, and only a fingerprint of the result is stored in the incident. This way needs the Reliability module.

Conclude closes a hypothesis as supported, refuted or inconclusive with a justification. Every conclusion needs attached evidence: supporting evidence for supported, refuting evidence for refuted, any evidence for inconclusive. A conclusion is final, and only a signed-in person can make it: a call made with an API token is refused with 403 human_conclusion_required, so an agent can narrow the cause down but the decision stays with a person.

Every change is checked against the incident version you saw. If someone changed the incident while your window was open, Mockarty refuses the change (409 revision_conflict), reloads the list and asks you to try again.

cURL

# Read the incident first: its "revision" goes into expectedRevision.
curl "$MOCKARTY_BASE_URL/api/v1/missions/incidents/$INCIDENT_ID?projectId=$PRODUCT_ID" \
  -H "Authorization: Bearer $MOCKARTY_API_KEY"

curl -X POST "$MOCKARTY_BASE_URL/api/v1/missions/incidents/$INCIDENT_ID/hypotheses" \
  -H "Authorization: Bearer $MOCKARTY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"projectId":"'"$PRODUCT_ID"'","code":"db_timeout","reason":"p99 latency tripled after the release",
       "idempotencyKey":"inc-42-db-timeout","expectedRevision":3}'

GET /api/v1/missions/incidents?projectId=… lists a product’s incidents newest first; pass the returned nextCursor as cursor for the next page. POST /api/v1/missions/incidents with projectId, missionId and an optional deploymentRunId opens one. POST /api/v1/missions/incidents/{incidentId}/hypotheses/{hypothesisId}/evidence attaches evidence: either sourceRef and kind (log, metric, trace, event or report), or source, queryKind and optional expression and windowMinutes; relation is supports, refutes or context. A source that is not configured returns 409 observability_unavailable. A configured source that fails to return valid evidence returns 503 observability_query_failed; check the source and retry. Repeating a request with the same idempotencyKey returns the first answer instead of adding a duplicate.

If an incident record fails an integrity check, the API returns 500 incident_integrity_error. Contact an administrator; changing the request cannot repair the record.

Metrics and budgets

The Metrics tab shows the totals for the period: how many missions finished and failed, the success rate, the tokens spent, how much of the budget is taken, plus spend by kind of work and by model. The Prefix cache tile shows how much of the missions’ input came from the provider’s prompt cache — cached input costs a fraction of the full price, so this is the share of spend the cache saved; zero means the cache is not being hit. The same number for one specific run is on its mission card. Only mission-attributed spend counts here — the workspace’s chat usage is not mixed in.

Open one mission and choose Economics for its exact immutable-ledger slice: model calls and tokens, cached input, physical runner/tool resources, events that do not yet have a price and the amount charged in each currency. The tab keeps the selected budget caps beside the actual consumption, but never mixes them. If accounting storage cannot be read, the card says that actual consumption is unavailable instead of presenting zero spend.

Budget utilisation is computed over unfinished work — it answers “will we hit the cap”, not “what did we spend”.

The same breakdowns are available as dashboard data sources — in the widget builder they sit under the Autonomous Missions category, so autonomous-work metrics can live on a shared dashboard next to everything else.


The same from an agent

Everything a person does on the page is available to an AI agent through Mockarty’s tools:

  • missions_list — the list and the counters in one call; productId=none together with the “waiting for data” state is the Triage queue;
  • mission_get — the mission card with its chain;
  • mission_start — start one;
  • mission_control — pause, resume, cancel, answer the mission’s question, and decide an approval gate; for pause, resume, cancel or an approval decision, use a stable idempotency_key when an agent may retry a lost response;
  • mission_step — retry or skip one node, with the same optional idempotency_key receipt;
  • mission_control_resolve — after direct executor inspection, safely resolve an ambiguous or interrupted receipt as applied or refused without replaying the original command;
  • attention_queue — what needs a person right now, in one call: missions waiting for an answer or an approval, running missions with anomalies, and missions failed in the last 24 hours — newest first. The same queue a person sees at the top of the missions page, so a human and an agent triage the same list;
  • incidents_list / incident_get / incident_open / incident_propose_hypothesis / incident_attach_evidence — the deployment incident journal: read incidents, open one on a mission’s deployment, propose a cause and back it with evidence. Concluding a hypothesis is left to a person, so there is no tool for it;
  • mission_review — the course check: is the mission doing what it was asked to;
  • mission_bind — bind to a product, optionally remembering the rule;
  • qa_cadence_set / qa_cadence_list — a product’s schedules, including their prompt;
  • capabilities_list — one catalogue of everything this instance can do: mission components, agent personas and tool profiles, runner engines, plugins and WASM extensions. Each versioned descriptor includes its exact source, trust and isolation state, side-effect and data-boundary policy, resource/cost envelope, schemas, executor/health binding, and your availability with a precise blocked reason. The common planner uses this authority automatically; operators and extension authors can inspect it to understand why a capability was selected or blocked. The same list is available as GET /api/v1/capabilities and as ListCapabilities / list_capabilities / listCapabilities in the Go, Python and Java SDKs;
  • mission_settings_effective — which autonomy settings were selected and why: each row includes the value, the layer that set it (mission → product → namespace → instance → builtin), and runtimeApplied. Guard policy and run window are frozen when a mission starts; the builtin run window is 480 minutes. Journal retention is different: it is a live namespace/instance operator policy consumed by the cleanup scheduler, so it is not frozen into one mission and product/mission overrides do not apply. Configure it in Autonomous Missions → Limits and settings → Mission evidence retention, leave either field blank to inherit, or automate it with mockarty-cli autonomy settings get|set and the Go/Python/Java NamespaceSettings SDK. Timeline summaries and terminal mission history use the event horizon; detailed payloads use the shorter payload horizon and are clamped so they never outlive their event. Active legal holds always win. Startup environment variables MOCKARTY_MISSION_EVENT_RETENTION_DAYS and MOCKARTY_MISSION_PAYLOAD_RETENTION_DAYS remain the instance builtin. The same effective answer is available as GET /api/v1/missions/settings/effective. A storage read failure returns 503 with code mission_settings_unavailable; cleanup fails closed rather than guessing a policy.

mission_start and mission_bind accept only a product that currently exists in the same namespace. An unknown or foreign product returns 404. If product authority cannot be read, they return retryable 503 product_authority_unavailable; retry the unchanged request. A repeated mission_start with the same originRef returns the already committed mission even if its former product was removed after the first response was lost.

Automating namespace safety and retention

The settings endpoint still replaces the older autonomy, budget and knowledge fields, so read the current document before changing those fields. The run wall and retention fields are additive: omitting one preserves its override; explicit null clears it and inherits. The run wall accepts 1..20160 minutes and defaults to 480 minutes when no layer overrides it. Retention accepts whole days from 1 to 3650, and detailed payload retention cannot exceed timeline retention. Requests are limited to 256 KiB and at most 128 knowledge references (kind up to 64 bytes, value up to 4096 bytes); incomplete rows and negative budgets are rejected instead of being silently changed. The CLI and SDK clear methods send null only when you explicitly request inheritance.

GET /api/v1/autotester/settings returns a strong ETag. Send it as If-Match when saving so a stale editor cannot overwrite a newer policy; the UI, CLI, and SDK helpers do this automatically. A stale write returns 412: reload and review the newer values. Every PUT also accepts Idempotency-Key; reuse the same key and exact body after a lost response. A different body under the same key returns 409. If the response is 503 with Policy-Applied: true and Audit-Pending: true, the settings are already saved and their central audit record is durably queued — retrying with the same key is safe.

mockarty-cli autonomy settings get --namespace engineering
mockarty-cli autonomy settings set --namespace engineering --run-window-minutes 90 --event-days 365 --payload-days 30 --request-id safety-change-2026-08-25
mockarty-cli autonomy settings set --namespace engineering --inherit-payload
mockarty-cli autonomy settings set --namespace engineering --inherit-run-window
current, _ := client.NamespaceSettings().GetAutonomySettings(ctx)
events, payloads, window := 365, 30, 90
current.JournalEventRetentionDays = &events
current.JournalPayloadRetentionDays = &payloads
current.RunWindowMinutes = &window
saved, err := client.NamespaceSettings().SaveAutonomySettingsWithOptions(ctx, current,
    mockarty.AutonomySettingsSaveOptions{RequestID: "retention-change-2026-08-23"})
// To intentionally clear a non-zero budget, also set ReplaceDefaultBudget: true.
// Explicitly inherit the instance payload policy:
saved, err = client.NamespaceSettings().ClearAutonomyRetention(ctx, false, true)
saved, err = client.NamespaceSettings().ClearAutonomyRunWindow(ctx, mockarty.AutonomySettingsSaveOptions{})
current = client.namespace_settings.get_autonomy_settings()
current.update(journalEventRetentionDays=365, journalPayloadRetentionDays=30, runWindowMinutes=90)
saved = client.namespace_settings.save_autonomy_settings(
    current, request_id="retention-change-2026-08-23")
# Explicitly inherit the instance payload policy:
saved = client.namespace_settings.clear_autonomy_retention(clear_payload=True)
saved = client.namespace_settings.clear_autonomy_run_window()
AutonomyNamespaceSettings current = client.namespaceSettings().getAutonomySettings();
current.journalEventRetentionDays(365).journalPayloadRetentionDays(30).runWindowMinutes(90);
AutonomyNamespaceSettings saved = client.namespaceSettings()
    .saveAutonomySettings(current, "retention-change-2026-08-23");
// Explicitly inherit the instance payload policy:
saved = client.namespaceSettings().clearAutonomyRetention(false, true);
saved = client.namespaceSettings().clearAutonomyRunWindow(null);

Common questions

A mission is “queued”. Open the card first. A delivery warning means the executor is temporarily unavailable and Mockarty is retrying automatically; it includes the next-attempt time. If there is no such warning, check that the required component is licensed and its service is running. A newly started mission with no executor is rejected immediately.

A scheduled run did not start. Look at that schedule’s latest mission: if the previous one is still open (waiting for approval, for instance), a new one is deliberately not created. Approve or close the previous one.

A mission landed in the wrong product. Press Bind and pick the right one. If such messages keep arriving, remove the stored rule from the wrong product’s intake settings and add it to the right one.


  • Autonomous Tester — how the autonomous tester works and what its autonomy levels mean.
  • Autonomous Coder — the runner fleet, the job queue and a product’s delivery settings.
  • Dashboards — where to put autonomous-work metrics.
  • Quality Verdicts — what a run’s verdict contains, and how a person withdraws or overrides one.