Docs Jury Panel

Jury Panel

Some things a test run must decide cannot be measured. Whether a page’s hero
“clearly states the offer”, whether a mission actually achieved the goal it was
given — no assertion settles those. Mockarty can put such questions to a panel of
language models that read the evidence the run gathered and vote.

The panel is configured in Settings → Jury. Nothing is enabled by default.

What the jury does — and what it must never do

The jury only ever fills a gap. If a deterministic check already settled
something, that result stands: a measurement is a fact, and no vote overrules a
fact. Concretely:

  • It judges only goals no engine could evaluate. Anything machine-checked is
    left alone.
  • It cannot turn a failure into a success. For a quality verdict its reading
    is recorded as advisory; for a mission it may downgrade a run that claimed
    success but not rehabilitate one that genuinely failed.
  • Disagreement makes the result claim less, not more. A split panel lowers
    the reported confidence of a verdict, and never overturns a recorded mission
    outcome.
  • A juror that cannot decide abstains. If nobody could answer, the goal stays
    reported as “declared, not evaluated” — which is the honest outcome, not a
    failure.

Seats, not profiles

A panel is a list of LLM profiles with a number of seats each. Seats are what
decide the jury:

  • Several seats on one profile ask the same model repeatedly. Where the
    evidence is clear the answers agree; where it is genuinely ambiguous they
    split — which is exactly the signal you want.
  • Several profiles additionally catch a single model’s bias.

Both are useful and both cost money: every seat is one model call per judged
goal
. The total is shown against a cap while you edit, and the server clamps
anything above it, so a panel can never quietly become a large bill.

Focus prompts: the same model, a different lens

Each member may carry an optional focus prompt — a stored prompt template
that tells that member what to weigh: “read this as an accessibility reviewer”,
“care about wording precision”, “think like a security engineer”.

This is the cheapest useful diversity a panel can have. The same model asked
twice answers roughly the same way; asked once as an accessibility reviewer and
once as a security reviewer, it produces two genuinely different readings.

A focus prompt adds a lens — it does not replace the rules. The instruction
that makes a vote trustworthy (judge only the evidence, abstain when unsure,
answer in a fixed shape) is stated before the focus text and restated after it,
so the last word always belongs to the contract. A focus prompt cannot instruct a
juror to always agree, to stop abstaining, or to answer in prose.

Because the lens is part of who a juror is, the same profile with two different
focus prompts counts as two members, not one — and each carries its own
seats.

If a focus prompt is deleted or cannot be read, that member votes under the base
rubric instead. Losing a lens costs some nuance; silently dropping the seat would
shrink the jury you configured.

Instance default and namespace overrides

Two levels:

  • Instance default — set by an administrator, used by every namespace that
    has no panel of its own.
  • Namespace override — replaces the default entirely for that namespace.

A namespace override replaces rather than merges. When you pay per seat you
need to look at one screen and see exactly the jury that will sit, and a panel
assembled from two levels is impossible to reason about.

The settings screen states which of the two you are looking at. If it says the
namespace is using the instance default, editing and saving creates an override;
“Drop override & inherit” removes it and returns to the default.

If neither level is configured, no jury runs anywhere.

Choosing profiles

Only enabled LLM profiles can sit on a panel — a disabled one would silently
produce a smaller jury than configured. If a profile is deleted after the panel
was saved, its row stays visible as unavailable rather than disappearing, so the
panel never shrinks without you noticing. Its seats are simply skipped on the
next run.

Where the reading appears

  • Quality verdicts — a judged criterion is marked as an advisory evaluation,
    and the verdict records which models judged it along with the prompt and rubric
    version, so a result stays reproducible and auditable.
  • Autonomous missions — the completion report carries whether the goal was
    met, whether the panel was split, and the vote count. A mission downgraded by
    the jury says so explicitly.