LLM API Mocks
When you build an application on top of an LLM provider (OpenAI, Azure OpenAI,
Anthropic, Ollama, OpenRouter, DeepSeek, …) you need to test it without calling
the real model: no cost, no rate limits, no flaky latency, and fully
deterministic answers in CI. A LLM API mock makes Mockarty answer like a
chat‑completion endpoint — including realistic token‑by‑token streaming — so
your app’s LLM client talks to Mockarty instead of the provider.
It is an ordinary HTTP mock with a special response mode: you match the
provider’s endpoint (e.g. POST /v1/chat/completions) and Mockarty returns a
provider‑shaped reply. Because it is a normal mock, everything else you already
know works on top of it — request conditions (answer differently per prompt
or model), dynamic content (echo the prompt, inject Faker data), priorities,
and namespaces.
What it produces
Mockarty looks at the stream flag in the incoming request body:
streamabsent orfalse→ a single JSON reply (chat.completionfor
OpenAI, amessageobject for Anthropic) with ausagetoken block.stream: true→ atext/event-streamthat emits the reply
token‑by‑token, exactly like the real provider (OpenAIchat.completion.chunk
deltas ending withdata: [DONE]; Anthropic’smessage_start→
content_block_delta→message_stopevent sequence).
Two provider formats are supported: openai (the de‑facto standard — also
covers Azure OpenAI, Ollama, OpenRouter, DeepSeek and most OpenAI‑compatible
clients) and anthropic (the Messages API).
Create one in the Constructor
- Open the Constructor, keep the protocol on HTTP, and set the route (e.g.
/v1/chat/completions) and method. - Under Response Body Type, click LLM.
- Pick the provider (OpenAI‑compatible or Anthropic), optionally set the model,
and write the reply content. Optionally set the finish reason and the stream
chunk size / delay. - Save. The mock now answers like a chat‑completion endpoint.
The same fields are available over the API as the llmResponse block below — use
whichever fits (the UI for humans, the API for scripts and AI agents).
Configure an LLM response
An LLM mock is a mock whose response carries an llmResponse block:
| Field | Meaning | Default |
|---|---|---|
provider |
openai or anthropic |
openai |
model |
Model name echoed back in the reply | the request’s model |
content |
The assistant reply text. Supports dynamic values (see below) | — (required) |
finishReason |
OpenAI finish_reason / Anthropic stop_reason |
stop / end_turn |
promptTokens |
Reported prompt tokens | auto‑estimated |
completionTokens |
Reported completion tokens | auto‑estimated |
chunkChars |
Characters per streamed delta (streaming only) | 4 (≈ one token) |
chunkDelayMs |
Delay before each streamed delta, in ms | 0 |
Dynamic content
content runs through the same templating engine as any mock payload, so the
reply can depend on the request:
- Echo the prompt / a request field — set
contentto a JsonPath into the
request body, e.g.$.req.modelor$.req.messages[-1].content. - Generate fake data — set
contentto a Faker expression, e.g.
$.fake.Sentenceor$.fake.FirstName.
(See the JsonPath Guide for the full expression syntax.)
Example
Create an OpenAI‑style streaming mock (adjust localhost:5770 to your server):
curl -X POST http://localhost:5770/api/v1/mocks \
-H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "chat-mock",
"namespace": "sandbox",
"http": { "route": "/v1/chat/completions", "httpMethod": "POST" },
"response": {
"statusCode": 200,
"llmResponse": {
"provider": "openai",
"content": "Hello! This is a mocked assistant reply.",
"chunkChars": 4,
"chunkDelayMs": 20
}
}
}'
Point your LLM client’s base URL at http://localhost:5770/stubs/sandbox and
call it as usual:
# Non-streaming → single JSON reply
curl -X POST http://localhost:5770/stubs/sandbox/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o","messages":[{"role":"user","content":"hi"}]}'
# Streaming → token-by-token text/event-stream
curl -N -X POST http://localhost:5770/stubs/sandbox/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o","stream":true,"messages":[{"role":"user","content":"hi"}]}'
The streaming call returns the reply one delta at a time:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hell"}}]}
…
data: [DONE]
Different answers per prompt
Because an LLM mock is a normal mock, add request conditions to return a
different reply depending on the request body (for example, a different content
when model equals gpt-4o, or when the prompt contains a keyword). Create one
mock per case on the same route; Mockarty picks the matching one by its
conditions and priority — the same way it does for any HTTP mock.
Resilience testing
To test how your application handles a misbehaving provider, give the mock a
non‑200 statusCode (e.g. 429 for rate‑limit) or a large chunkDelayMs to
simulate a slow stream. Combine LLM mocks with Mockarty’s chaos features to
inject latency and faults on the same endpoint.
Tool / function calls
To test function-calling (tool-use) flows, return tool calls instead of (or
alongside) text. In the Constructor open Tool calls and enter a JSON array;
over the API set llmResponse.toolCalls:
"llmResponse": {
"provider": "openai",
"content": "",
"toolCalls": [{ "name": "get_weather", "arguments": { "city": "NYC" } }]
}
The mock returns the provider’s tool-call shape — OpenAI message.tool_calls with
finish_reason: "tool_calls" (arguments serialised to a JSON string); Anthropic a
tool_use content block with stop_reason: "tool_use". The reply text may be
empty for a tool-only response.
Multi-turn conversations
A chat client resends the whole conversation in messages on every call, so a
mock can answer the latest turn and keep state across turns:
- Reply to the latest message — set
contentto$.req.lastUserMessage
(the last user turn),$.req.lastMessage,$.req.lastAssistantMessage, or
$.req.systemPrompt.$.req.turnCountis the number of messages so far. (Use
these instead of$.req.messages[N].content, which works only with a fixed
positive index.) - Branch per turn / per content — add request conditions on
$.req.messages(e.g. match when the conversation contains a keyword) and
create one mock per case; Mockarty picks the matching one. - Remember across turns — add an Extract rule that writes a request value
into a Chain or Global store (e.g.cStore.lastQuestion = $.req.lastUserMessage).
The next request reads it back with$.cS.lastQuestion— so turn N’s reply can
reference what happened on turn N‑1. (See Store Systems.)
Example — an assistant that always answers the latest user message:
"llmResponse": { "provider": "openai", "content": "You said: $.req.lastUserMessage" }
See also
- Scripted Responses — full programmatic control of a reply
- JsonPath Guide — dynamic content expressions
- Store Systems — stateful multi‑turn scenarios