Docs LLM API Mocks

LLM API Mocks

When you build an application on top of an LLM provider (OpenAI, Azure OpenAI,
Anthropic, Ollama, OpenRouter, DeepSeek, …) you need to test it without calling
the real model: no cost, no rate limits, no flaky latency, and fully
deterministic answers in CI. A LLM API mock makes Mockarty answer like a
chat‑completion endpoint — including realistic token‑by‑token streaming — so
your app’s LLM client talks to Mockarty instead of the provider.

It is an ordinary HTTP mock with a special response mode: you match the
provider’s endpoint (e.g. POST /v1/chat/completions) and Mockarty returns a
provider‑shaped reply. Because it is a normal mock, everything else you already
know works on top of it — request conditions (answer differently per prompt
or model), dynamic content (echo the prompt, inject Faker data), priorities,
and namespaces.

What it produces

Mockarty looks at the stream flag in the incoming request body:

  • stream absent or false → a single JSON reply (chat.completion for
    OpenAI, a message object for Anthropic) with a usage token block.
  • stream: true → a text/event-stream that emits the reply
    token‑by‑token, exactly like the real provider (OpenAI chat.completion.chunk
    deltas ending with data: [DONE]; Anthropic’s message_start →
    content_block_delta → message_stop event sequence).

Two provider formats are supported: openai (the de‑facto standard — also
covers Azure OpenAI, Ollama, OpenRouter, DeepSeek and most OpenAI‑compatible
clients) and anthropic (the Messages API).

Create one in the Constructor

  1. Open the Constructor, keep the protocol on HTTP, and set the route (e.g.
    /v1/chat/completions) and method.
  2. Under Response Body Type, click LLM.
  3. Pick the provider (OpenAI‑compatible or Anthropic), optionally set the model,
    and write the reply content. Optionally set the finish reason and the stream
    chunk size / delay.
  4. Save. The mock now answers like a chat‑completion endpoint.

The same fields are available over the API as the llmResponse block below — use
whichever fits (the UI for humans, the API for scripts and AI agents).

Configure an LLM response

An LLM mock is a mock whose response carries an llmResponse block:

Field Meaning Default
provider openai or anthropic openai
model Model name echoed back in the reply the request’s model
content The assistant reply text. Supports dynamic values (see below) — (required)
finishReason OpenAI finish_reason / Anthropic stop_reason stop / end_turn
promptTokens Reported prompt tokens auto‑estimated
completionTokens Reported completion tokens auto‑estimated
chunkChars Characters per streamed delta (streaming only) 4 (≈ one token)
chunkDelayMs Delay before each streamed delta, in ms 0

Dynamic content

content runs through the same templating engine as any mock payload, so the
reply can depend on the request:

  • Echo the prompt / a request field — set content to a JsonPath into the
    request body, e.g. $.req.model or $.req.messages[-1].content.
  • Generate fake data — set content to a Faker expression, e.g.
    $.fake.Sentence or $.fake.FirstName.

(See the JsonPath Guide for the full expression syntax.)

Example

Create an OpenAI‑style streaming mock (adjust localhost:5770 to your server):

curl -X POST http://localhost:5770/api/v1/mocks \
  -H "Authorization: Bearer $MOCKARTY_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "chat-mock",
    "namespace": "sandbox",
    "http": { "route": "/v1/chat/completions", "httpMethod": "POST" },
    "response": {
      "statusCode": 200,
      "llmResponse": {
        "provider": "openai",
        "content": "Hello! This is a mocked assistant reply.",
        "chunkChars": 4,
        "chunkDelayMs": 20
      }
    }
  }'

Point your LLM client’s base URL at http://localhost:5770/stubs/sandbox and
call it as usual:

# Non-streaming → single JSON reply
curl -X POST http://localhost:5770/stubs/sandbox/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o","messages":[{"role":"user","content":"hi"}]}'

# Streaming → token-by-token text/event-stream
curl -N -X POST http://localhost:5770/stubs/sandbox/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o","stream":true,"messages":[{"role":"user","content":"hi"}]}'

The streaming call returns the reply one delta at a time:

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hell"}}]}

…

data: [DONE]

Different answers per prompt

Because an LLM mock is a normal mock, add request conditions to return a
different reply depending on the request body (for example, a different content
when model equals gpt-4o, or when the prompt contains a keyword). Create one
mock per case on the same route; Mockarty picks the matching one by its
conditions and priority — the same way it does for any HTTP mock.

Resilience testing

To test how your application handles a misbehaving provider, give the mock a
non‑200 statusCode (e.g. 429 for rate‑limit) or a large chunkDelayMs to
simulate a slow stream. Combine LLM mocks with Mockarty’s chaos features to
inject latency and faults on the same endpoint.

Tool / function calls

To test function-calling (tool-use) flows, return tool calls instead of (or
alongside) text. In the Constructor open Tool calls and enter a JSON array;
over the API set llmResponse.toolCalls:

"llmResponse": {
  "provider": "openai",
  "content": "",
  "toolCalls": [{ "name": "get_weather", "arguments": { "city": "NYC" } }]
}

The mock returns the provider’s tool-call shape — OpenAI message.tool_calls with
finish_reason: "tool_calls" (arguments serialised to a JSON string); Anthropic a
tool_use content block with stop_reason: "tool_use". The reply text may be
empty for a tool-only response.

Multi-turn conversations

A chat client resends the whole conversation in messages on every call, so a
mock can answer the latest turn and keep state across turns:

  • Reply to the latest message — set content to $.req.lastUserMessage
    (the last user turn), $.req.lastMessage, $.req.lastAssistantMessage, or
    $.req.systemPrompt. $.req.turnCount is the number of messages so far. (Use
    these instead of $.req.messages[N].content, which works only with a fixed
    positive index.)
  • Branch per turn / per content — add request conditions on
    $.req.messages (e.g. match when the conversation contains a keyword) and
    create one mock per case; Mockarty picks the matching one.
  • Remember across turns — add an Extract rule that writes a request value
    into a Chain or Global store (e.g. cStore.lastQuestion = $.req.lastUserMessage).
    The next request reads it back with $.cS.lastQuestion — so turn N’s reply can
    reference what happened on turn N‑1. (See Store Systems.)

Example — an assistant that always answers the latest user message:

"llmResponse": { "provider": "openai", "content": "You said: $.req.lastUserMessage" }

See also