Autonomous Tester
The Autonomous Tester is Mockarty’s autonomous testing agent. You hand it a goal in plain language — “regression-test the checkout API against its spec and fuzz the payment endpoint” — and it orients itself in the product under test, plans an approach, and (depending on how much freedom you give it) runs the testing for you. It records every step, creates artifacts (mocks, collections, test cases, findings), and reports back. A person can watch it work and take over at any moment.
About URLs in examples: All examples use
localhost:5770as the default Mockarty address. If your instance runs on a remote server, replacelocalhost:5770with its actual address (e.g.https://mockarty.company.comorhttp://192.168.1.50:5770). See Tips & Useful Features for details.
The Autonomous Tester is available on plans that include autonomous agents: in Mockarty Cloud that is Team and Enterprise (Pro and Pro+ keep the AI buttons, the chat and the MCP server). If the page or endpoints are not available on your instance, the feature is not part of your current plan.
What it is
A mission is one testing goal you give the agent. From the goal, the agent:
- Orients itself — reads the knowledge you have fed it (uploaded docs, specs, pull-request descriptions) so it understands what the product does before touching it.
- Plans — decides which kinds of testing are relevant (contract, API, fuzz, UI, security, load, test cases).
- Acts — depending on the autonomy level, either waits for your approval or runs the plan and reports results.
You supervise the whole thing from the Autonomous Missions cockpit in the sidebar. Its live mission list and detail card keep the step flow, evidence, budget, settings, and controls in one place.
Autonomy levels
When you launch a mission you choose how much initiative the agent takes:
| Level | What it does | When to use it |
|---|---|---|
| recon | Read-only investigation. The agent observes and reports, but makes no changes. | You want a safe first look — an inventory of what could be tested — with zero risk of side effects. |
| propose (default) | The agent plans, then stops and waits for you to approve before running anything. This is the human-in-the-loop default. | Supervised runs. You want to see the plan and approve it before the agent acts. |
| auto | Fully autonomous end-to-end. The agent plans and runs with no approval gate. | Unattended runs — CI pipelines, agent-driven automation, overnight runs where no human is watching. Choose this deliberately. |
propose is the default because a human gate is the safer choice. Pass auto on purpose when you know no one will be at the keyboard.
Launching a mission
From the UI
- Open Autonomous Missions in the sidebar.
- Click Launch mission.
- Type the Goal — describe what the agent should test in plain language.
- Pick the Autonomy level (
propose,auto, orrecon). - Optionally set a Token budget — a hard cap on total tokens the mission may spend. Leaving it empty does not mean unlimited: a generous safety ceiling still applies, so a mission can never run away with your tokens. Set the budget explicitly when you want a tighter — or a larger — bound.
- Click submit (or press Ctrl+Enter in the goal field).
The new mission appears in the list and starts working (or, in propose mode, plans and waits for your approval).
From the API
Send an authenticated POST to /api/v1/autotester/intents with your goal. The response gives you a missionId you can use to follow progress.
Send the credential in Authorization: Bearer … or X-API-Key; credentials in
the URL query string are not accepted. If an API or integration token has an
allowed_actions restriction, it must include write to start a mission. A
read-only token receives 403 before the request body is parsed and creates no
mission.
The token must be bound to a namespace. A mission runs unattended under the namespace its token belongs to, so a token without one is refused with 401 — and adding ?namespace= to the URL does not change that. Create the token with a namespace:
curl -X POST http://localhost:5770/api/v1/auth/tokens \
-H "Authorization: Bearer $SESSION_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"autonomous-agent","namespace":"my-team","allowed_actions":["read","write"]}'
If your product has no web page at /, declare at least one endpoint. The agent
checks that the product is really there before spending anything, and a root that
answers 404 looks like nothing is deployed. Declaring a real handler proves it is:
"contextRefs": [
{"kind": "endpoint", "value": "GET /api/orders"},
{"kind": "spec", "value": "https://api.example.com/openapi.json"}
]
curl -X POST http://localhost:5770/api/v1/autotester/intents \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"goal":"Regression-test the checkout API against its OpenAPI spec and fuzz the payment endpoint","autonomy":"propose","budget":{"tokens_total":500000}}'
Response:
{"missionId":"m_9f2c…","status":"accepted"}
For a repeatable acceptance campaign, the request may also include
rmaAssessment with a case ID, the campaign corpus SHA-256 digest, and ordered
oracle:1, oracle:2, … then forbidden:1, … criteria. Set a non-empty
traceId (up to 128 UTF-8 bytes) in the same request. Mockarty validates the
declaration and binds its
exact criteria, case, corpus and trace to the mission’s immutable brief before
execution. The server does not approve the supplied corpus digest. This records
what the campaign asked to check; it does not mean
the mission passed those checks. A semantic outcome remains unavailable until
a separate server-side assessment has been recorded.
The official SDKs and CLI expose the same submit-and-supervise flow:
Go
accepted, err := client.AutonomousMissions().Submit(ctx, mockarty.AutonomousMissionSubmitRequest{
Goal: "Regression-test checkout", Autonomy: "auto",
Budget: mockarty.AutonomousMissionBudgetHint{TokensTotal: 100000},
})
page, err := client.AutonomousMissions().List(ctx, "active", 50)
mission, err := client.AutonomousMissions().Get(ctx, accepted.MissionID)
flow, err := client.AutonomousMissions().GetFlow(ctx, accepted.MissionID)
Python
request = AutonomousMissionSubmitRequest(
goal="Regression-test checkout", autonomy="auto",
budget=AutonomousMissionBudgetHint(tokens_total=100_000),
)
accepted = client.autonomous_missions.submit(request)
page = client.autonomous_missions.list(status="active", limit=50)
mission = client.autonomous_missions.get(accepted.mission_id)
flow = client.autonomous_missions.get_flow(accepted.mission_id)
Java
var request = new AutonomousMissionSubmitRequest()
.goal("Regression-test checkout").autonomy("auto").budget(100000, 0, 0);
var accepted = client.autonomousMissions().submit(request);
var page = client.autonomousMissions().list("active", 50);
var mission = client.autonomousMissions().get(accepted.getMissionId());
var flow = client.autonomousMissions().getFlow(accepted.getMissionId());
CLI
mockarty-cli autonomy missions submit --goal "Regression-test checkout" \
--autonomy auto --tokens-total 100000
mockarty-cli autonomy missions list --status active --limit 50
mockarty-cli autonomy missions get "$MISSION_ID"
mockarty-cli autonomy missions flow "$MISSION_ID"
Repeat the flow read while the mission is non-terminal. Terminal statuses are
done and failed; awaiting_approval and awaiting_info require operator input.
The body accepts:
goal(required) — plain-language description of what to test.autonomy—recon,propose, orauto. Omit it and the agent usespropose.budget— optional non-negative token and USD limits. Zero leaves that limit at the server default; negative,NaN, and infinite values are rejected.
Keep the missionId to poll the mission (see Watching progress).
For AI agents (MCP)
An AI agent can start a mission with the submit_testing_task MCP tool (goal + optional autonomy/budget/context hints), and follow it with autotester_mission_flow to inspect the mission end-to-end in a single call. See AI Features and the MCP Marketplace for how tools are exposed to agents.
Approving a plan
In propose mode the mission plans and then waits. On the mission detail page you’ll see the Proposed plan with two buttons:
- Approve & run — the agent executes the approved plan.
- Reject — the mission is stopped without running the plan.
Over the API:
# Approve the proposed plan and start execution
curl -X POST http://localhost:5770/api/v1/autotester/missions/$MISSION_ID/approve \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
# Reject the plan (mission is stopped)
curl -X POST http://localhost:5770/api/v1/autotester/missions/$MISSION_ID/reject \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
Both return {"ok":true} on success.
Answering the agent
Sometimes a mission needs a fact it can’t find on its own — a base URL, a credential name, which environment to hit. When that happens the mission pauses and shows an Awaiting info banner with the question. Type your answer in the box on the mission page and press Send answer (or Ctrl+Enter), and the mission resumes.
Over the API, the mission’s details include the pending question (awaitingQuestion) and the id you must echo back (awaitingRequestId). Post your answer:
curl -X POST http://localhost:5770/api/v1/autotester/missions/$MISSION_ID/info-response \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"requestId":"<awaitingRequestId>","available":true,"answers":{"baseUrl":"https://staging.example.com"}}'
requestId(required) — theawaitingRequestIdfrom the mission.available—trueif you’re providing the answer,falseif the information isn’t available.answers— a map of the facts the agent asked for.
Returns {"ok":true}, and the mission continues.
Watching progress
The mission list shows each mission’s goal and status. Statuses you’ll see include active, paused, awaiting_approval (waiting for you to approve a plan), awaiting_info (waiting for your answer), done, and failed.
Select a mission to see:
- Autonomy and step count.
- Budget — tokens spent against the cap (see Budgets).
- Mission flow — an ordered list of executed steps (each with a summary, timestamp, token cost, and a failed flag if it went wrong).
- Created — the artifacts the mission produced (mocks, collections, test cases, findings), with links.
- Controls — Pause, Resume, and Stop.
The page updates live as the mission works.
Over the API you can poll:
# Full mission flow in one call: mission + steps + artifacts
curl http://localhost:5770/api/v1/autotester/missions/$MISSION_ID/flow \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
Other read endpoints: GET /api/v1/autotester/missions (list), GET /api/v1/autotester/missions/{id} (one mission), .../timeline (steps), .../artifacts (created entities). Control endpoints mirror the buttons: POST .../pause, .../resume, .../stop. The legacy .../stop route is a compatibility alias for durable Mission cancellation when the run has already been admitted to the common Mission control plane. Reuse an optional Idempotency-Key header after a lost response; HTTP 202 means cancellation was accepted and is still being delivered, not that every remote executor has already stopped.
For an externally submitted mission with declared RMA acceptance criteria, read
GET /api/v1/autotester/missions/{id}/assessment with the same mission-read
permission. It returns the server’s terminal result with one outcome for each
declared criterion. pass means every criterion has matching evidence and
independent review; fail means at least one criterion has evidenced failure;
could_not_verify means the available evidence was insufficient. A missing or
foreign result returns 404. A claimed result or one whose mission evidence
changed returns 409; a temporary source failure returns 503. This read does
not start a new mission or repeat model calls.
If an administrator has selected journal-only history and Mockarty cannot prove that projection is complete, the flow, timeline, and agent-tasks reads return 503 Service Unavailable instead of silently showing stale legacy history. Retry after the administrator repairs or rolls back the journal cutover.
Feeding it knowledge
The agent tests better when it understands the product first. You can feed it context it will read before planning:
- Documents and specs — upload docs, specs, or notes into the product-context knowledge base.
- Pull requests / merge requests — point the agent at a GitHub PR or GitLab MR URL. It fetches the title, description, and diff (redacting secrets) and stores them so the agent can find the rationale behind a recently merged feature.
For AI agents, this is exposed as the product_context_ingest_pr MCP tool (ingest a PR/MR by URL) and product_context_search (search the knowledge base). The same knowledge base powers Mockarty’s Knowledge Base (RAG) features.
Ingest a PR over the API:
curl -X POST http://localhost:5770/api/v1/autotester/context/docs/ingest-pr \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url":"https://github.com/owner/repo/pull/123"}'
For a private repository, add a token field with a personal access token. Search the knowledge base:
curl "http://localhost:5770/api/v1/autotester/context/search?query=payment%20API%20validation%20rules&k=5" \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
Searching across the three corpora
The agent reads knowledge from three sources, and they are now distinguishable in every answer and selectable in every search:
| Corpus | What it holds | Search tool |
|---|---|---|
product_context |
Documents, specs and ingested PR/MR descriptions that describe the product | product_context_search |
experience |
What earlier runs found out — pitfalls, mission lessons, product facts, defect root causes | experience_search |
security_kb |
Vulnerability taxonomy, remediation playbooks and uploaded security documents | security_kb_search |
Every result carries a corpus field, so a caller can tell a specification from a previous run’s observation without guessing. All three search tools accept an optional corpus argument that narrows the answer to one corpus; omit it to search everything, exactly as before.
# Only run-experience records, not product documents
curl "http://localhost:5770/api/v1/autotester/context/knowledge/search?query=payment%20retry&corpus=experience" \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
# Only documents, not earlier runs' findings
curl "http://localhost:5770/api/v1/autotester/context/search?query=payment%20retry&corpus=product_context" \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
A corpus value outside the three names (security_kb, product_context, experience) is rejected with a message naming the known ones — a filter is either applied or refused, never silently ignored. A record whose corpus cannot be determined answers unknown rather than a guessed name.
Reusing experience from earlier runs
Product context answers “what do the documents say?” Run experience answers “what did earlier work actually discover?” The experience corpus stores four kinds of reusable observation: pitfall, mission_lesson, product_fact, and defect_root_cause. Every entry includes a source and provenance. Treat it as evidence to verify, not as a confirmed specification.
AI agents use experience_search before rediscovering an address, authentication rule, recurring failure, or earlier fix. They use experience_record after learning something concrete. A record is created as a review candidate and cannot influence later missions until an authorized reviewer publishes it. Repeating the exact same evidence is idempotent; it does not replace a different entry. Every entry answers with its corpus (experience), and the corpus is assigned by the platform from the entry kind — a request body cannot declare its entry to belong to another corpus.
cURL
curl "http://localhost:5770/api/v1/autotester/context/knowledge/search?query=payment%20retry&kinds=pitfall&k=5" \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
curl -X POST http://localhost:5770/api/v1/autotester/context/knowledge \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"kind":"pitfall","text":"Payment retries require an idempotency key","source":"autonomous run m-42, turn 17"}'
Go SDK
items, err := client.Experience().Search(ctx, mockarty.ExperienceSearchRequest{
Query: "payment retry", Kinds: []string{mockarty.ExperienceKindPitfall}, Limit: 5,
})
Python SDK
items = client.experience.search(query="payment retry", kinds=["pitfall"], limit=5)
Java SDK
ExperienceSearchResponse items = client.experience()
.search("payment retry", List.of(ExperienceApi.KIND_PITFALL), null, 5);
Knowledge from outside your installation
Some tools let an agent leave your installation to find an answer: it can fetch
a page to read it, or call a third-party MCP server. Whether that is acceptable
is your decision, not the agent’s — an isolated contour usually wants it closed,
and a connected one often wants it narrowed to a few known hosts.
MOCKARTY_EXTERNAL_KNOWLEDGE is that decision:
| Value | What an agent may consult outside your installation |
|---|---|
unset, or open |
any public address (the default) |
allowlist |
only the hosts you name in MOCKARTY_EXTERNAL_KNOWLEDGE_HOSTS |
off |
nothing — the agent answers from your workspace, the product documentation and project memory |
MOCKARTY_EXTERNAL_KNOWLEDGE_HOSTS is a comma-separated list. An entry is
either an exact host (docs.example.com) or, starting with a dot,
that domain and everything under it (.example.com).
Anything else is rejected and treated as off, so a typo can never widen what
an agent may reach. With off the outside-reaching tools are not offered to the
agent at all, and it plans without them instead of trying and failing.
This setting is about KNOWLEDGE. It does not touch the systems you are testing:
requests to a target you named — API calls, load tests, scans — keep working
exactly as before.
Product Quality Profile
A product under test can be tracked as a long-lived entity instead of a one-off run. A product accumulates the facts the tester should reuse (base URLs, spec references, examples, known issues) and a quality trend over time, so each run starts from what the last run learned — regression, deltas, and a rolling risk level — rather than from scratch.
Namespace credentials for platform services
These endpoints use the namespace carried by the credential unless you select
one explicitly with ?namespace=... or X-Mockarty-Namespace. A 403 saying
that the credential is scoped to another namespace describes the credential,
not whether the requested namespace exists. Use an API key scoped to that
namespace. If one platform service must work across customer namespaces, an
administrator can create an admin-scope integration token in Admin Panel →
Integrations (or with POST /api/v1/integrations) and use its one-time mki_...
token. Admin scope is cross-namespace authority; do not use it for a single
customer integration, and store the token as a secret.
Create or update a product (the key is a stable slug, unique per namespace — the same key updates in place, never a duplicate):
curl -X POST http://localhost:5770/api/v1/qa/products \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"key":"payments-api","name":"Payments API",
"baseUrls":["https://staging.example.com"],
"specRefs":["https://staging.example.com/openapi.json"]}'
Attach a reusable fact (an example, an invariant, a known issue). Re-adding the same sourceRef updates it in place, so ingesting the same PR twice doesn’t pile up:
curl -X POST http://localhost:5770/api/v1/qa/products/<id>/facts \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"kind":"known_issue","body":"Checkout 500s under load","sourceRef":"issue:42"}'
Read the quality trend — the current score/risk plus the per-dimension snapshots each run appends:
curl http://localhost:5770/api/v1/qa/products/<id>/quality \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
Read everything already known about a product in one call — before starting a rework round. The response carries the testing plan for this product in plain prose, the open and resolved defects, per-area coverage, and how the last run ended:
curl http://localhost:5770/api/v1/qa/products/<key-or-id>/knowledge \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN"
Why one call rather than four: whoever is fixing is starting a round, and if orienting takes several requests they make none of them and re-discover what earlier runs already found.
What is useful to a fixer specifically:
strategy.brief— how we decided to check this product and why, including a section on what the plan does not cover;strategy.history— earlier versions of the plan (number, who authored it, when and why it changed). It answers the question the current plan cannot: has our understanding of the product moved since the run whose defects you are fixing;openDefects[].key— the same key a run’s defect carries, so a defect in a report and a row in the product’s record are the same thing;openDefects[].where— the address: an endpoint, a screen or a step. The difference between searching and fixing;lastRun.sufficiency— whether the CHECKING was enough. Separate from whether the product passed: a run can find no defects and still be insufficient.notRecheckableDefects[]— defects whose surface this build can no longer exercise on this product (the plan excludes the area, or the battery cannot run here). They are deliberately NOT inopenDefects: nobody can close them by any action, so they are not live breaks. Each carriesnotRecheckableReason; one becomes open again if the surface becomes testable and the finding reproduces.
For AI agents, the same surface is exposed as MCP tools: qa_product_list, qa_product_get, qa_product_upsert, qa_product_quality, qa_product_facts_add, qa_product_delete, and product_knowledge.
Continuous testing (cadences)
A product can run on a schedule, around the clock — a nightly security scan, a weekly load test, an hourly smoke run. Each cadence fires an autonomous mission on its cron, with goals derived from its kind and a bounded budget envelope so a 24/7 schedule never overspends.
curl -X POST http://localhost:5770/api/v1/qa/products/<id>/cadences \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"kind":"security","cronExpr":"0 2 * * *","enabled":true}'
kind is one of security, load, regression, contract, smoke, or full (the complete matrix). cronExpr accepts a standard five-field cron (0 2 * * * = 02:00 daily) or a descriptor (@daily, @every 1h). The response includes the computed nextFireAt. For AI agents: qa_cadence_list, qa_cadence_set, qa_cadence_delete.
Reacting to your team’s work
When your team uses Mockarty’s tracker, wiki, test cases, and chat, the autonomous tester watches those signals and reacts on its own — accumulating each change into the product-context knowledge base, and (where you’ve configured a rule) starting a testing mission from it. Anti-spam throttling means a burst of activity never becomes a burst of missions. This turns the tester into a continuous teammate that keeps up with the product as it changes, not just something you launch by hand.
Budgets
You can cap how much a mission is allowed to spend so it stops before it overspends. When launching a mission (for example through the submit_testing_task tool or the API), you can set:
- tokens_total — a hard cap on total tokens for the mission.
- tokens_per_day — a per-day token cap.
- usd_cap — an optional spend cap in USD, counted from what the mission is actually charged at your price list. Each step is also held within the money that is left, so a single long step cannot run far past the cap. When the cap is reached the mission pauses with a budget notice; raise the cap and resume to continue.
The mission page shows tokens spent against the cap so you can see how much budget is left at a glance.
Connectors & triggers
Missions don’t have to be started by hand. Open Autonomous Missions → Triggers (/ui/missions?tab=triggers) and use Configure intake to connect external systems so that missions start automatically — for example from an issue tracker, a chat bot, a message bus, or a signed webhook. Configure the connector once, and incoming events turn into testing missions the agent picks up on its own. See External Trackers & Tools and CI Triggers for related integration surfaces.
Signed webhook (HTTP intake)
An inbound HTTP Webhook connector gives you a public endpoint an external system can call to start missions. When you create one, the UI shows the endpoint URL (it is also visible in the connector list, with a copy button). The sender must:
POSTthe intent JSON (same shape as/api/v1/autotester/intents— at minimum{"goal":"..."}) to/api/v1/intake/<connector-id>.- Send the target namespace in the
X-Namespaceheader. - Sign the raw request body with the connector’s HMAC secret — hex-encoded HMAC-SHA256 — and send it in the
X-Intake-Signatureheader.
BODY='{"goal":"Smoke-test the orders API","autonomy":"auto"}'
SIG=$(printf '%s' "$BODY" | openssl dgst -sha256 -hmac "$INTAKE_SECRET" -hex | sed 's/^.* //')
curl -X POST http://localhost:5770/api/v1/intake/$CONNECTOR_ID \
-H "X-Namespace: default" \
-H "X-Intake-Signature: $SIG" \
-H "Content-Type: application/json" \
-d "$BODY"
A valid signature returns 202 Accepted and the mission appears in the Autonomous Missions cockpit. A missing or wrong signature is rejected with 401.
Related
- AI Features — the AI agent chat and built-in tools.
- MCP Marketplace — how tools are exposed to AI agents.
- Knowledge Base (RAG) — the product-context store.
- Security Agent — autonomous security scanning.