Runner as an MCP server — give your AI agent a browser
Any AI agent that speaks MCP — Claude Code, Cursor, your own SDK agent — can
drive the Mockarty runner directly: browse pages, read them, click, type, take
screenshots, and control an attached Android phone.
The browser tools need no Mockarty server, account, or licence key. Download
the runner, start it in MCP mode, point your agent at it — that surface is free
to use anywhere, as an alternative to running a full browser automation stack.
Device control and the rest of Mockarty’s testing mechanics are part of the
platform: they appear only once this runner has shaken hands with a Mockarty
server (COORDINATOR_URL + API_TOKEN, on a build carrying the Mockarty
licence key). An unpaired runner does not list them at all — you always see
exactly the tools that will work.
Why an agent wants this
A full desktop browser costs 50–150 MB per instance and starts slowly. The
runner uses a lightweight headless engine: hundreds of parallel sessions fit on
a small machine, and pages come back as text or markdown, which is what a
model actually reads — no screenshot round-trip needed to understand a page.
Start it
# stdio — the transport local MCP clients launch themselves
mockarty-runner mcp
# or over the network (token required)
MOCKARTY_RUNNER_MCP_TOKEN=your-secret mockarty-runner mcp --http 127.0.0.1:9800
Connect an MCP client. For Claude Code, add to .mcp.json:
{
"mcpServers": {
"mockarty-runner": {
"command": "mockarty-runner",
"args": ["mcp"]
}
}
}
Over HTTP instead:
{
"mcpServers": {
"mockarty-runner": {
"type": "http",
"url": "http://127.0.0.1:9800/mcp",
"headers": { "X-API-Key": "your-secret" }
}
}
}
The tools
Call runner_info first — it reports what this machine can do (browser engine,
attached devices, platform), so the agent never guesses.
| Tool | What it does |
|---|---|
runner_info |
Capabilities of this runner — call first |
| Seeing | |
browser_snapshot |
The page’s controls as ref=N role "name" — act by ref, no CSS guessing. A huge page is capped and says so; browser_find still reaches a control past the cap |
browser_read |
The page’s content: text, markdown, outline, links. A very large page is capped (200 KB / 500 links) and the reply says so — scope with a selector or browser_find for the rest |
browser_find |
Find elements by visible name → refs |
browser_screenshot |
PNG of the page (when pixels matter). Returns a download link by default (output="link") so a big image never floods your context; pass output="base64" only to feed the pixels to a vision step |
browser_pdf |
Save the page as a PDF → download link (or output="base64"). Print-layout checks, archiving. Chromium engine only — not advertised on a light-engine runner |
| Navigating | |
browser_open / browser_goto |
Open a session / navigate it. browser_open accepts storageState (from browser_storage_state) to start already logged-in |
browser_back / browser_forward / browser_reload |
History |
browser_wait_for |
Wait for text to appear/disappear, a selector, or a delay |
| Interacting | |
browser_act |
click / fill / press — by ref or CSS selector |
browser_fill_form |
Fill a whole form in one call (+ optional submit selector or submitRef snapshot ref) — log in / register in a single tool call |
browser_hover |
Hover (menus, tooltips) |
browser_select |
Pick option(s) in a <select> |
browser_check |
Tick / untick a checkbox or radio |
browser_upload |
Attach files to a file input |
browser_drag |
Drag one element onto another — each end by selector (from/to) or snapshot ref (fromRef/toRef) |
browser_key |
Send a key to the page (Escape, Enter, Control+A) |
browser_resize |
Change the viewport (responsive checks) |
| Tabs & diagnostics | |
browser_tabs |
list / new / select / close |
browser_capture |
Start recording console + network |
browser_console |
Captured console output (JS errors) |
browser_network |
Captured requests (method + URL) |
| Session | |
browser_eval |
Evaluate a JS expression (result capped ~200 KB — return a small value, not outerHTML) |
browser_sessions |
List open sessions |
browser_storage_state |
Export cookies + localStorage (log in once, reuse) |
browser_close |
Close a session |
| Verifying | |
browser_case |
Run a whole test CASE in one call — a JSON array of goto/back/forward/reload/act/select/check/key/hover/drag/resize/upload/fill_form/wait/assert steps → a PASS/FAIL report; stops at the first failing step. Browser case testing beyond step-by-step driving |
browser_assert |
PASS/FAIL check: text (contains or regex), visibility, value, count, url, title, attribute, checked, enabled |
browser_dialog |
Accept or dismiss the next alert / confirm / prompt |
browser_element_screenshot |
Screenshot one element (download link by default; output="base64" to inline) |
browser_visual_diff |
PASS/FAIL against a baseline image YOU pass — the fraction of changed pixels; the highlighted diff comes back as a link on fail |
| Auditing (out of the box — Playwright MCP has neither) | |
browser_perf |
Performance report: Lighthouse-style web vitals from the Performance API — TTFB, DOM interactive / content-loaded, load, First Contentful Paint, best-effort Largest Contentful Paint, request count, transfer bytes |
browser_a11y |
Accessibility audit: title/lang, heading outline, landmark count, and the WCAG smells to flag — images without alt, controls with no label, buttons/links with no accessible name |
Mobile tools (paired runners)
When the runner is connected to a Mockarty server AND an Android device is
attached (adb on PATH, USB debugging on), the mobile surface appears —
everything an agent needs to drive a phone, with no Appium server to install:
| Tool | What it does |
|---|---|
device_list |
Attached devices |
device_info |
Screen size, density, Android version, model |
device_source |
The screen’s UI hierarchy as ref=N Class "text" @x,y — act by ref. A dense screen is capped and says so; narrow with filter or device_find |
device_find |
Find elements by text / content-desc / resource id |
device_tap |
Tap by ref (preferred) or by x,y |
device_text |
Type into the focused field |
device_swipe |
Swipe by direction=up/down/left/right or explicit coordinates |
device_key |
back, home, enter, recent, delete, volume… |
device_screenshot |
PNG of the screen |
device_app |
launch / terminate / clear / current / list |
device_wait_for |
Wait until text appears or disappears |
device_assert |
PASS/FAIL: text_contains, text_absent, element_visible, element_absent, app_is |
device_case |
Run a whole ON-DEVICE test CASE in one call — a JSON array of app/tap/text/swipe/key/wait/assert steps → a PASS/FAIL report; stops at the first failing step (the mobile analog of browser_case) |
runner_info tells you which mode you are in and, in public mode, how to unlock
the rest.
A mobile flow
device_app {"deviceId":"…","action":"launch","package":"com.example.app"}
device_wait_for{"deviceId":"…","text":"Sign in"}
device_source {"deviceId":"…"} → ref=5 Button "Sign in" @541,750
device_tap {"deviceId":"…","ref":5}
device_assert {"deviceId":"…","kind":"text_contains","expected":"Welcome"}
Tapping by ref — not by pixels — is what makes a mobile flow survive a
different screen size, and what makes the transcript readable later.
The workflow that works
browser_open {"url": "https://example.com"} → sessionId
browser_snapshot {"sessionId": "s1"} → ref=2 textbox "Username", ref=1 button "Sign in"
browser_act {"sessionId": "s1", "ref": 2, "kind": "fill", "value": "demo"}
browser_act {"sessionId": "s1", "ref": 1} → click
browser_wait_for {"sessionId": "s1", "text": "Welcome"} → wait for the async result
browser_read {"sessionId": "s1", "mode": "markdown"} → confirm
browser_close {"sessionId": "s1"}
Reuse one session for a whole flow. browser_snapshot tells you what you can
click; browser_read tells you what the page says. Acting by ref is more
reliable than a hand-written CSS selector and survives re-renders. After anything
async, browser_wait_for instead of polling with reads.
Once the flow is settled, wrap it in browser_case to run it as one repeatable
test with a single PASS/FAIL verdict — the way CI wants it:
browser_case {"url": "https://app/login", "steps": "[
{\"do\":\"fill_form\",\"fields\":[{\"selector\":\"#user\",\"value\":\"demo\"}],\"submit\":\"button[type=submit]\"},
{\"do\":\"wait\",\"text\":\"Welcome\"},
{\"do\":\"assert\",\"kind\":\"text_contains\",\"expected\":\"Welcome, demo\"}
]"}
→ CASE PASSED (3/3 steps) (or CASE FAILED at step N — stops there)
To run a case behind a login in one call, pass storageState (the JSON a
previous browser_storage_state returned) alongside url — the fresh session
starts already authenticated, so the whole case runs logged-in without a login
step. Requires the Chromium engine (the light engine cannot restore a saved
state).
Step kinds a case can run: goto (navigate), back/forward/reload (navigation history), act (click/fill/press by selector
or ref), select (pick values in a <select>), check (tick a checkbox —
checked defaults to true; pass false to untick), key (a keyboard key —
Enter/Escape/Tab/ArrowDown), hover (move the pointer over a
selector/ref — reveals hover menus), drag (drag a source selector/ref onto
a target to/toRef), resize (set the viewport width×height — responsive
testing at any resolution), upload (attach files on the runner to a file
<input> — over the network transport this needs RUNNER_MCP_UPLOAD_DIR),
fill_form (a whole form, optional
submit selector or submitRef), wait (for text/selector/gone, optional ms), and assert
(text_contains/text_matches/text_absent/visible/hidden/value/count/url_contains/title_contains/attribute/checked/enabled/disabled). That covers a real form —
dropdowns and checkboxes included — in a single call.
Debugging a page that misbehaves: browser_capture → reproduce → browser_console
and browser_network.
What it will not do
- Only
httpandhttps.file://and other schemes are refused, so a
page read can never turn into a local file read. - Targets are validated. Cloud metadata endpoints and malformed hosts are
rejected. If you deliberately test such an address, set
MOCKARTY_VIEW_ALLOW_ANY_TARGET=1— one switch, and it lifts the same check
for every browser Mockarty drives, including UI tests and interactive
sessions, not just this one. - Platform tools stay locked without the handshake. Devices and the
backend-testing mechanics are Mockarty features; the public surface is the
browser. - The HTTP transport requires a token. These tools drive a real browser on
your machine that can reach whatever your machine can reach; the runner
refuses to start the network transport without one. - Treat the token like shell access. A holder can point the browser at any
http(s) target your machine can reach, read what any page renders, and attach
any local file the runner process can read (browser_upload) to a page. Bind
--httpto loopback (127.0.0.1) or a trusted network, and hand the token
only to agents you trust the way you would trust a shell on this machine.
Options
| Flag / variable | Meaning |
|---|---|
--http <addr> |
Serve streamable HTTP instead of stdio |
--max-sessions N |
Browser session cap (default 100, LRU-evicted) |
--session-ttl D |
Idle session lifetime (default 10m) |
MOCKARTY_RUNNER_MCP_TOKEN |
Token for the HTTP transport |
RUNNER_MCP_ENGINE |
Browser backend: chromium (default) or light |
RUNNER_BROWSER_AUTO_INSTALL |
0 to skip the one-time browser download (air-gapped hosts) |
RUNNER_ADB_BINARY |
Path to adb when it is not on PATH |
VIEWCORE_ACT_TIMEOUT_MS |
How long browser_act / browser_fill_form wait for an element before failing (default 15000). Lower it for even snappier feedback on wrong selectors; raise it for pages whose elements appear slowly. |
RUNNER_MCP_UPLOAD_DIR |
Over the HTTP transport, browser_upload may only read files from this directory (a remote client must not read arbitrary server files). Unset over HTTP = browser_upload refused. Ignored over stdio, where the client is the local user. |
Which browser engine
By default the server drives a full Chromium. It opens https:// pages on
any machine, and the browser is downloaded automatically the first time you open
a page — nothing to install by hand.
Set RUNNER_MCP_ENGINE=light to use the lightweight engine instead: it
holds far more sessions in parallel on the same memory, which is ideal for
large-scale scripted runs. It needs the host’s certificate store to open
https:// pages, so use it on a server where that store is present rather than
on a laptop. The light engine’s servo render core resizes the window per
screenshot, so per-shot viewport changes (browser_resize before a shot,
different viewport values) are honoured on both engines — a test can
screenshot the same session at several resolutions on either.
Air-gapped: a runner with the engines inside it
The default build downloads a browser engine the first time it needs one. In an
isolated environment there is nothing to download from, so the runner can be
built self-contained: the lightweight engines are baked into the binary and
unpacked on first use. One file, zero network, works behind an air gap.
GOOS=linux GOARCH=amd64 bash scripts/fetch-engines.sh # fetch the engines once
go build -tags embedengines -o mockarty-runner ./cmd/mockarty-runner
The result is a single binary (~270 MB, the engines are inside) that runs
browser_* tools with no downloads at all. Without the tag you get the lean
build (~70 MB) that provisions engines on demand — the right choice when the
machine has internet.
Real browsers from Playwright (Chromium/Firefox/WebKit) stay a separate,
pluggable concern: they are never embedded. Point the runner at a pre-installed
browser instead, or pin the built-in engines to files you placed yourself with
RUNNER_LIGHTPANDA_PATH / RUNNER_SERVO_PATH.
Then what?
The same runner can join a Mockarty grid (COORDINATOR_URL + API_TOKEN) and
execute UI, load, fuzz and mobile tests dispatched from the platform, with
reports, history and a device grid. See Browser runner
and Runner labels.