Quality Verdicts
Every autonomous quality run — a mission that tests a product on its own — ends
with a verdict: a document that says what was checked, what held,
what did not, and, just as importantly, what was not checked and why. This
page explains what a verdict contains, how it is accepted, and the two things
a person can do about one afterwards: withdraw it or override it.
What a verdict is
A verdict belongs to one run and is identified by the run’s task id. It carries:
- Outcome —
pass,failorcould_not_verify. The last one is a real
answer, not a shrug: the run could not reach a conclusion (an unreachable
product, an exhausted budget) and says so instead of guessing. - Surfaces — every area the run was asked to examine, each with
ran
(whether it was exercised at all), its coverage, depth, confidence and the
findings it produced. A surface withran: falsewas never touched; its
silence is not a pass. Coverage such as18/29means 18 checks passed out of
29 executed; failed checks appear in findings, while checks that did not run
appear in uncovered gaps. - Findings — each with a severity, a reproduction and a
dedup_keythat
matches the product-knowledge defects, so a new problem can be told from one
an earlier run already reported. - Uncovered gaps — what was not checked, with a
remedythat says what
would actually help: fix the product or its environment, retry (a transient
or budget limit), nothing (verification unavailable), or a person’s judgement. - The jury record — the rubric and prompt versions, the models that voted,
how many criteria were settled, how many were disputed and how many jurors
abstained, plus a calibrated confidence. - Acceptance binding — whether this verdict is still the current one for
its run, or has been withdrawn.
The conformance block reports a result for each criterion sent in
acceptance[]. Describing requirements in product_context or naming targets
in conformance_target does not create individual criteria; without
acceptance[], conformance is absent.
A criterion whose target is an endpoint ({"kind":"endpoint","ref":"POST /api/orders"})
is judged only by checks of that endpoint. If nothing checked it — for example
because writes are switched off for the run — the criterion is reported as not
evaluated, with the reason and what to declare so it can be measured. Findings
on other endpoints never fail it.
Acceptance is automatic
A verdict is accepted the moment it is published. There is no approval step:
in an automated flow an external system is waiting for the answer and nobody
is there to press a button, so holding the verdict for a person who is not
coming would only mean it is never processed. What a person gets instead is
an after-the-fact decision, for the runs they chose to supervise.
Reading verdicts
# The latest verdicts of your namespace — one headline row each
curl -H "X-API-Key: $MOCKARTY_API_KEY" \
"http://localhost:5770/api/v1/quality/verdicts?limit=20"
# One verdict in full
curl -H "X-API-Key: $MOCKARTY_API_KEY" \
"http://localhost:5770/api/v1/quality/verdicts/<taskId>"
The list answers {"verdicts": [...], "namespace": "...", "count": N}; each
row carries taskId, outcome, qualityScore, notEvaluated, current,
bindingStatus, the jury headline and confidenceLowerBound. When the run’s
client declared the build it tested, the row also carries
acceptedArtifactDigest — the artifact digest derived from that commit, so you
can compare it with the deployment view below without opening the verdict.
Bodies are not included in the list — read one verdict for the full document,
which comes back under verdict. limit accepts 1–200 (default 20);
namespace defaults to the caller’s.
The full verdict also includes a deployment block: the same comparison the
dedicated endpoint below performs, computed on the spot.
Is this verdict what is deployed?
One request compares the build a verdict vouches for with the candidate currently
deployed to an environment. Both sides are digests derived from a source
commit, so the comparison is exact and needs no manual work:
curl -H "X-API-Key: $MOCKARTY_API_KEY" \
"http://localhost:5770/api/v1/quality/verdicts/<taskId>/deployment?environment=staging"
The answer carries state, acceptedDigest (from the verdict), deployedDigest
(from the accepted deployment run), and deploymentRunId:
exact— the deployed candidate is the build the verdict vouches for;different— something else is deployed; treat the verdict as history, not as
a description of the live environment;unknown— the latest deployment is still applying or its rollback is not
verified, no accepted deployment is known, the verdict declared no build, or
deployments are not journaled on this installation.
environment is optional: without it the latest deployment run with a possible
effect in the namespace determines the comparison. When that run is accepted,
the response reports its environment.
Withdrawing a verdict (invalidate)
Withdrawing records that the verdict is no longer to be trusted. The run’s
binding switches to invalidated, with the reason and who withdrew it. It does
not un-send what an external consumer already received.
curl -X POST -H "X-API-Key: $MOCKARTY_API_KEY" -H "Content-Type: application/json" \
-d '{"reason": "the stand was mid-deploy during this run"}' \
"http://localhost:5770/api/v1/quality/verdicts/<taskId>/invalidate"
The reason is optional but recorded; the answer names the new status and
repeats the warning that the consumer was not notified.
Overriding a verdict (override)
Overriding replaces the conclusion with a human decision. It creates a new
revision of the verdict: the machine revision is never edited, the evidence
(what was measured) is carried forward unchanged, and the reason travels with
the decision. The accepted-verdict projection advances in the same transaction,
so the verdict authority and what the rest of the platform reads can never
disagree about one run.
curl -X POST -H "X-API-Key: $MOCKARTY_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: review-2026-09-21-run-42" \
-d '{"reason": "the failing check targets a flag that is off in production", "outcome": "pass", "expectedRevision": 1}' \
"http://localhost:5770/api/v1/quality/verdicts/<taskId>/override"
reasonis required — an override without a reason is not a decision
anyone can review.outcomeis optional (pass,fail,could_not_verify) and defaults to
the verdict’s current outcome, so a reason can be attached without
changing the conclusion.expectedRevisionis optional and defaults to the current revision; pass
it when you read the verdict earlier and want the override refused with
409if somebody changed it in between.Idempotency-Keymakes a retry safe; when omitted a deterministic key is
derived. The answer is{verdictId, namespace, revision, outcome, reason, replayed}—replayed: truemeans this exact override had already been
recorded and nothing new was written.
Who may do what
Reading verdicts needs an authenticated session or API token with access to
the namespace. Withdrawing and overriding need write access to the
namespace the verdict belongs to; any other caller gets 403. Every
withdrawal and override is written to the audit trail with the actor and the
reason. Verdicts belong to the Autonomous Missions module and follow its licence.
For agents
Reading and overriding are also MCP tools, so an AI agent can read a run’s
result before acting on it and record a supervised decision:
quality_verdicts_list— the headline rows (limit), including the
acceptedArtifactDigestof every run that declared a build.quality_verdict_get— the full document for ataskId, with the
deploymentcomparison inside; readsurfaces[].rananduncovered[]
before the score.aqc_override_verdict—taskIdandreasonrequired,outcomeand
expectedRevisionoptional, same semantics as the REST call.
Withdrawing (invalidate) is a REST-only operation.
Related pages
- Autonomous Missions
- Autonomous Agent
- Admin Setup — the decision journal