Docs Runner as an MCP server (give your agent a browser)

Runner as an MCP server — give your AI agent a browser

Any AI agent that speaks MCP — Claude Code, Cursor, your own SDK agent — can
drive the Mockarty runner directly: browse pages, read them, click, type, take
screenshots, and control an attached Android phone.

The browser tools need no Mockarty server, account, or licence key. Download
the runner, start it in MCP mode, point your agent at it — that surface is free
to use anywhere, as an alternative to running a full browser automation stack.

Device control and the rest of Mockarty’s testing mechanics are part of the
platform: they appear only once this runner has shaken hands with a Mockarty
server (COORDINATOR_URL + API_TOKEN, on a build carrying the Mockarty
licence key). An unpaired runner does not list them at all — you always see
exactly the tools that will work.

Why an agent wants this

A full desktop browser costs 50–150 MB per instance and starts slowly. The
runner uses a lightweight headless engine: hundreds of parallel sessions fit on
a small machine, and pages come back as text or markdown, which is what a
model actually reads — no screenshot round-trip needed to understand a page.

Start it

# stdio — the transport local MCP clients launch themselves
mockarty-runner mcp

# or over the network (token required)
MOCKARTY_RUNNER_MCP_TOKEN=your-secret mockarty-runner mcp --http 127.0.0.1:9800

Connect an MCP client. For Claude Code, add to .mcp.json:

{
  "mcpServers": {
    "mockarty-runner": {
      "command": "mockarty-runner",
      "args": ["mcp"]
    }
  }
}

Over HTTP instead:

{
  "mcpServers": {
    "mockarty-runner": {
      "type": "http",
      "url": "http://127.0.0.1:9800/mcp",
      "headers": { "X-API-Key": "your-secret" }
    }
  }
}

The tools

Call runner_info first — it reports what this machine can do (browser engine,
attached devices, platform), so the agent never guesses.

Tool What it does
runner_info Capabilities of this runner — call first
Seeing
browser_snapshot The page’s controls as ref=N role "name" — act by ref, no CSS guessing. A huge page is capped and says so; browser_find still reaches a control past the cap
browser_read The page’s content: text, markdown, outline, links. A very large page is capped (200 KB / 500 links) and the reply says so — scope with a selector or browser_find for the rest
browser_find Find elements by visible name → refs
browser_screenshot PNG of the page (when pixels matter). Returns a download link by default (output="link") so a big image never floods your context; pass output="base64" only to feed the pixels to a vision step
browser_pdf Save the page as a PDF → download link (or output="base64"). Print-layout checks, archiving. Chromium engine only — not advertised on a light-engine runner
Navigating
browser_open / browser_goto Open a session / navigate it. browser_open accepts storageState (from browser_storage_state) to start already logged-in
browser_back / browser_forward / browser_reload History
browser_wait_for Wait for text to appear/disappear, a selector, or a delay
Interacting
browser_act click / fill / press — by ref or CSS selector
browser_fill_form Fill a whole form in one call (+ optional submit selector or submitRef snapshot ref) — log in / register in a single tool call
browser_hover Hover (menus, tooltips)
browser_select Pick option(s) in a <select>
browser_check Tick / untick a checkbox or radio
browser_upload Attach files to a file input
browser_drag Drag one element onto another — each end by selector (from/to) or snapshot ref (fromRef/toRef)
browser_key Send a key to the page (Escape, Enter, Control+A)
browser_resize Change the viewport (responsive checks)
Tabs & diagnostics
browser_tabs list / new / select / close
browser_capture Start recording console + network
browser_console Captured console output (JS errors)
browser_network Captured requests (method + URL)
Session
browser_eval Evaluate a JS expression (result capped ~200 KB — return a small value, not outerHTML)
browser_sessions List open sessions
browser_storage_state Export cookies + localStorage (log in once, reuse)
browser_close Close a session
Verifying
browser_case Run a whole test CASE in one call — a JSON array of goto/back/forward/reload/act/select/check/key/hover/drag/resize/upload/fill_form/wait/assert steps → a PASS/FAIL report; stops at the first failing step. Browser case testing beyond step-by-step driving
browser_assert PASS/FAIL check: text (contains or regex), visibility, value, count, url, title, attribute, checked, enabled
browser_dialog Accept or dismiss the next alert / confirm / prompt
browser_element_screenshot Screenshot one element (download link by default; output="base64" to inline)
browser_visual_diff PASS/FAIL against a baseline image YOU pass — the fraction of changed pixels; the highlighted diff comes back as a link on fail
Auditing (out of the box — Playwright MCP has neither)
browser_perf Performance report: Lighthouse-style web vitals from the Performance API — TTFB, DOM interactive / content-loaded, load, First Contentful Paint, best-effort Largest Contentful Paint, request count, transfer bytes
browser_a11y Accessibility audit: title/lang, heading outline, landmark count, and the WCAG smells to flag — images without alt, controls with no label, buttons/links with no accessible name

Mobile tools (paired runners)

When the runner is connected to a Mockarty server AND an Android device is
attached (adb on PATH, USB debugging on), the mobile surface appears —
everything an agent needs to drive a phone, with no Appium server to install:

Tool What it does
device_list Attached devices
device_info Screen size, density, Android version, model
device_source The screen’s UI hierarchy as ref=N Class "text" @x,y — act by ref. A dense screen is capped and says so; narrow with filter or device_find
device_find Find elements by text / content-desc / resource id
device_tap Tap by ref (preferred) or by x,y
device_text Type into the focused field
device_swipe Swipe by direction=up/down/left/right or explicit coordinates
device_key back, home, enter, recent, delete, volume…
device_screenshot PNG of the screen
device_app launch / terminate / clear / current / list
device_wait_for Wait until text appears or disappears
device_assert PASS/FAIL: text_contains, text_absent, element_visible, element_absent, app_is
device_case Run a whole ON-DEVICE test CASE in one call — a JSON array of app/tap/text/swipe/key/wait/assert steps → a PASS/FAIL report; stops at the first failing step (the mobile analog of browser_case)

runner_info tells you which mode you are in and, in public mode, how to unlock
the rest.

A mobile flow

device_app     {"deviceId":"…","action":"launch","package":"com.example.app"}
device_wait_for{"deviceId":"…","text":"Sign in"}
device_source  {"deviceId":"…"}              → ref=5 Button "Sign in" @541,750
device_tap     {"deviceId":"…","ref":5}
device_assert  {"deviceId":"…","kind":"text_contains","expected":"Welcome"}

Tapping by ref — not by pixels — is what makes a mobile flow survive a
different screen size, and what makes the transcript readable later.

The workflow that works

browser_open     {"url": "https://example.com"}          → sessionId
browser_snapshot {"sessionId": "s1"}                     → ref=2 textbox "Username", ref=1 button "Sign in"
browser_act      {"sessionId": "s1", "ref": 2, "kind": "fill", "value": "demo"}
browser_act      {"sessionId": "s1", "ref": 1}           → click
browser_wait_for {"sessionId": "s1", "text": "Welcome"}  → wait for the async result
browser_read     {"sessionId": "s1", "mode": "markdown"} → confirm
browser_close    {"sessionId": "s1"}

Reuse one session for a whole flow. browser_snapshot tells you what you can
click; browser_read tells you what the page says. Acting by ref is more
reliable than a hand-written CSS selector and survives re-renders. After anything
async, browser_wait_for instead of polling with reads.

Once the flow is settled, wrap it in browser_case to run it as one repeatable
test with a single PASS/FAIL verdict — the way CI wants it:

browser_case {"url": "https://app/login", "steps": "[
  {\"do\":\"fill_form\",\"fields\":[{\"selector\":\"#user\",\"value\":\"demo\"}],\"submit\":\"button[type=submit]\"},
  {\"do\":\"wait\",\"text\":\"Welcome\"},
  {\"do\":\"assert\",\"kind\":\"text_contains\",\"expected\":\"Welcome, demo\"}
]"}
→ CASE PASSED (3/3 steps)     (or CASE FAILED at step N — stops there)

To run a case behind a login in one call, pass storageState (the JSON a
previous browser_storage_state returned) alongside url — the fresh session
starts already authenticated, so the whole case runs logged-in without a login
step. Requires the Chromium engine (the light engine cannot restore a saved
state).

Step kinds a case can run: goto (navigate), back/forward/reload (navigation history), act (click/fill/press by selector
or ref), select (pick values in a <select>), check (tick a checkbox —
checked defaults to true; pass false to untick), key (a keyboard key —
Enter/Escape/Tab/ArrowDown), hover (move the pointer over a
selector/ref — reveals hover menus), drag (drag a source selector/ref onto
a target to/toRef), resize (set the viewport width×height — responsive
testing at any resolution), upload (attach files on the runner to a file
<input> — over the network transport this needs RUNNER_MCP_UPLOAD_DIR),
fill_form (a whole form, optional
submit selector or submitRef), wait (for text/selector/gone, optional ms), and assert
(text_contains/text_matches/text_absent/visible/hidden/value/count/url_contains/title_contains/attribute/checked/enabled/disabled). That covers a real form —
dropdowns and checkboxes included — in a single call.

Debugging a page that misbehaves: browser_capture → reproduce → browser_console
and browser_network.

What it will not do

  • Only http and https. file:// and other schemes are refused, so a
    page read can never turn into a local file read.
  • Targets are validated. Cloud metadata endpoints and malformed hosts are
    rejected. If you deliberately test such an address, set
    MOCKARTY_VIEW_ALLOW_ANY_TARGET=1 — one switch, and it lifts the same check
    for every browser Mockarty drives, including UI tests and interactive
    sessions, not just this one.
  • Platform tools stay locked without the handshake. Devices and the
    backend-testing mechanics are Mockarty features; the public surface is the
    browser.
  • The HTTP transport requires a token. These tools drive a real browser on
    your machine that can reach whatever your machine can reach; the runner
    refuses to start the network transport without one.
  • Treat the token like shell access. A holder can point the browser at any
    http(s) target your machine can reach, read what any page renders, and attach
    any local file the runner process can read (browser_upload) to a page. Bind
    --http to loopback (127.0.0.1) or a trusted network, and hand the token
    only to agents you trust the way you would trust a shell on this machine.

Options

Flag / variable Meaning
--http <addr> Serve streamable HTTP instead of stdio
--max-sessions N Browser session cap (default 100, LRU-evicted)
--session-ttl D Idle session lifetime (default 10m)
MOCKARTY_RUNNER_MCP_TOKEN Token for the HTTP transport
RUNNER_MCP_ENGINE Browser backend: chromium (default) or light
RUNNER_BROWSER_AUTO_INSTALL 0 to skip the one-time browser download (air-gapped hosts)
RUNNER_ADB_BINARY Path to adb when it is not on PATH
VIEWCORE_ACT_TIMEOUT_MS How long browser_act / browser_fill_form wait for an element before failing (default 15000). Lower it for even snappier feedback on wrong selectors; raise it for pages whose elements appear slowly.
RUNNER_MCP_UPLOAD_DIR Over the HTTP transport, browser_upload may only read files from this directory (a remote client must not read arbitrary server files). Unset over HTTP = browser_upload refused. Ignored over stdio, where the client is the local user.

Which browser engine

By default the server drives a full Chromium. It opens https:// pages on
any machine, and the browser is downloaded automatically the first time you open
a page — nothing to install by hand.

Set RUNNER_MCP_ENGINE=light to use the lightweight engine instead: it
holds far more sessions in parallel on the same memory, which is ideal for
large-scale scripted runs. It needs the host’s certificate store to open
https:// pages, so use it on a server where that store is present rather than
on a laptop. The light engine’s servo render core resizes the window per
screenshot, so per-shot viewport changes (browser_resize before a shot,
different viewport values) are honoured on both engines — a test can
screenshot the same session at several resolutions on either.

Air-gapped: a runner with the engines inside it

The default build downloads a browser engine the first time it needs one. In an
isolated environment there is nothing to download from, so the runner can be
built self-contained: the lightweight engines are baked into the binary and
unpacked on first use. One file, zero network, works behind an air gap.

GOOS=linux GOARCH=amd64 bash scripts/fetch-engines.sh   # fetch the engines once
go build -tags embedengines -o mockarty-runner ./cmd/mockarty-runner

The result is a single binary (~270 MB, the engines are inside) that runs
browser_* tools with no downloads at all. Without the tag you get the lean
build (~70 MB) that provisions engines on demand — the right choice when the
machine has internet.

Real browsers from Playwright (Chromium/Firefox/WebKit) stay a separate,
pluggable concern: they are never embedded. Point the runner at a pre-installed
browser instead, or pin the built-in engines to files you placed yourself with
RUNNER_LIGHTPANDA_PATH / RUNNER_SERVO_PATH.

Then what?

The same runner can join a Mockarty grid (COORDINATOR_URL + API_TOKEN) and
execute UI, load, fuzz and mobile tests dispatched from the platform, with
reports, history and a device grid. See Browser runner
and Runner labels.