Docs Bot Testing

Bot Testing

Bot Testing verifies a chat-bot the only way that proves it works — as a conversation. You script a dialog as ordered turns (a user stimulus and the reply the bot must give), run it against a mocked Bot API, and get a per-turn verdict transcript.

A single request can’t accept a bot: correctness lives in the exchange — command → reply, state carried across turns, inline-button round-trips. Bot Testing runs exactly that.

A scenario is a list of turns

Each turn has:

  • a stimulus — what the simulated user does: message (text), command (a /command), callback (an inline-button press), url (follow the link a URL button carries), or first_button (take the first choice the bot just offered);
  • an expectation — what the bot must do in response within a window: send_message, send_keyboard, edit_message, answer_callback, url_open (a url turn’s link opened), or none (the negative assertion — the bot must stay silent). Text is matched by textContains / textRegex; keyboards by buttonContains.

A turn passes when its expectation is met inside the reply window; the run passes when every turn passes.

Taking a button without knowing its label

Menu-driven bots label their buttons differently in every product, and those labels change with the copy. Writing them into a scenario makes it break on the next wording edit.

first_button takes the first button of the keyboard the bot showed on the previous turn, whichever keyboard it used — an inline button is pressed by its callback data, a URL button by following its link, a reply-keyboard button by sending its label back. You never write the label:

{"turns": [
  {"stimulus": {"kind": "command", "text": "/start"},
   "expectation": {"kind": "send_keyboard"}},
  {"stimulus": {"kind": "first_button"},
   "expectation": {"kind": "send_message", "textContains": "date"}}
]}

If the bot offered no buttons at that point, the turn fails and says so — that is a real finding about the flow, because it is exactly where a person gets stuck.

Pressing a URL button

A bot can offer a choice that opens a link instead of addressing the bot ({"text": "Docs", "url": "…"} on the inline keyboard). The url stimulus follows that link the way a chat client opens it, and the outcome of the open is the turn verdict:

{"turns": [
  {"stimulus": {"kind": "command", "text": "/start"},
   "expectation": {"kind": "send_keyboard"}},
  {"stimulus": {"kind": "first_button"},
   "expectation": {"kind": "url_open"}}
]}

The turn passes when the link answers successfully (recorded in the transcript as a url_open action), and fails with the reason when the landing is broken (HTTP 500), unreachable, or refused. A link on localhost or a private address is refused by default — the same outbound guard that protects every other outgoing call — unless the operator has enabled calls to private addresses (ALLOW_PROXY_TO_PRIVATE_IPS=true); the refusal message says so. This is what makes “does the bot’s link still work” an assertable step instead of a manual check.

In the scenario editor, choose User does → opens a link button and paste the link, or presses the first offered button to take whatever the bot showed; for a link press the Bot must field is set to open the link successfully for you.

The mocked Bot API

Bot Testing runs an emulated Telegram Bot-API server. Point your bot’s api_url at the session it hands you (every bot library supports a custom API URL — the standard self-hosted bot-api switch) and your bot behaves exactly as against the real server: it long-polls getUpdates and calls sendMessage / editMessageText / answerCallbackQuery. The suite injects the user stimuli as updates and reads the bot’s calls back as the turn’s observed actions.

memory is a built-in double for scripted scenarios that don’t need a running bot.

Bot Testing is multi-platform. The platform layer is a registry, so it is not tied to one messenger:

  • telegram-botapi — the emulated Telegram Bot-API described above.
  • generic — a neutral HTTP-bot protocol any framework can target: your bot polls GET <baseUrl>/updates for user stimuli and POSTs its replies to POST <baseUrl>/reply with {kind?, text, buttons?}. Use this for a custom bot or any messenger without a dedicated adapter.
  • memory — the scripted double.

A new concrete messenger (VK Teams, Max, Slack…) is an additive adapter — no change to scenarios, the API, or the UI.

The bot stand: point a real bot at a fake Telegram

A stand is a long-lived emulated Bot-API server you drive by hand — no scenario, no test plan. Use it when you are building the bot and just want to talk to it.

  1. Specialized Testing → Bot Testing → Bot stand → New stand.
  2. Copy the Bot API URL and the Bot token into your bot’s configuration (every library exposes an API-URL setting — the standard self-hosted Bot-API switch, e.g. base_url / api_url / a custom API endpoint).
  3. Start the bot. It works exactly as against the real server: it long-polls getUpdates, or registers a webhook with setWebhook.
  4. Send an update — a message, a /command, or a button press — and everything the bot does appears under What the bot did. Buttons the bot offers are clickable: one click sends the right press back (callback data for an inline button, the label for a reply keyboard).

A stand is idle-expiring (30 minutes by default, up to 24 hours via ttlSeconds) and belongs to the namespace that created it — nobody else can see or drive it.

Webhook bots

Bots that receive updates by webhook are tested the same way, with no code change.

  • If your bot calls setWebhook itself, nothing else is needed: the stand records the URL and starts POSTing updates there.
  • If the webhook is configured outside the bot, set it from the stand: Webhook → Webhook URL (+ an optional secret token), or POST /bot-suite/sessions/{id}/webhook.

Every delivery is a real HTTP POST to your bot carrying the Update JSON and the X-Telegram-Bot-Api-Secret-Token header, so a bot that validates the secret keeps validating it. The stand shows each delivery with its status code, attempt count and error, so “the update never arrived” is visible instead of guessed. A bot may also answer the webhook request with the method call itself ({"method": "sendMessage", ...}) — that reply is recorded as a bot action just like a direct API call.

Polling and webhook are mutually exclusive, exactly as on the real server: while a webhook is registered, getUpdates answers 409 Conflict; deleteWebhook returns the stand to long-polling.

Reaching a bot on a private address. A bot running on localhost or inside your container network is only reachable when the operator has enabled outbound calls to private addresses (ALLOW_PROXY_TO_PRIVATE_IPS=true). Without it the webhook URL is refused up front with a message that says so — outbound requests to internal addresses are blocked by default. Cloud-metadata addresses stay blocked in every configuration.

Supported Bot API methods

The stand answers every method: the ones below are understood (they become assertable actions with their text, caption and keyboard), and anything else answers ok: true and is recorded as a generic api_call, so an unsupported method never breaks your bot.

Area Methods
Identity getMe, getChat
Receiving updates getUpdates (long-poll, offset / limit / allowed_updates), setWebhook, deleteWebhook, getWebhookInfo
Messages sendMessage, editMessageText, editMessageCaption, editMessageReplyMarkup, deleteMessage
Media sendPhoto, sendDocument, sendVideo, sendAudio, sendVoice, sendAnimation, sendSticker, sendVideoNote, sendMediaGroup
Interaction answerCallbackQuery, answerInlineQuery, sendChatAction
Files getFile + downloading the file back from the returned file_path
Setup setMyCommands, getMyCommands, deleteMyCommands, setChatMenuButton, setMyDescription, setMyName, logOut, close

Both keyboard kinds are read: inline_keyboard (pressed by callback data) and keyboard (pressed by sending the label back). Requests are accepted as JSON, form-encoded, query-string or multipart upload — whatever your library sends.

Matching expectations follow the same vocabulary: send_message, send_keyboard, edit_message, answer_callback, send_photo, send_document, send_media, send_chat_action, answer_inline_query, delete_message, api_call, url_open (a url turn’s link opened), and none for the silence assertion.

In the UI

Bot Testing in the sidebar, with two tabs.

Scenarios — the middle pane is a chat: each turn shows the user bubble and the expected-bot bubble; click a turn to edit its stimulus and expectation on the right. Run annotates every bubble with its verdict and latency.

Bot stand — the live stand described above: create one, copy its URL and token into your bot, send updates, watch what the bot does, and manage its webhook.

Via the API

Run an inline scenario in one call. A telegram-botapi or generic run requires sessionId — the stand you created above and pointed the bot at. Without it the run is refused with offendingField: sessionId and a hint naming the stand endpoint, instead of being minted into a stand nobody is polling (the usual cause of an all-fail transcript). platform: "memory" is the scripted double and needs no stand.

The generic stand observes send_message and send_keyboard replies only — edit_message and answer_callback are telegram-specific. bot_stand_create accepts platform: "generic" for such bots (default is the Telegram emulation).

Run an inline scenario in one call:

curl -X POST "http://localhost:5770/api/v1/namespaces/my-namespace/bot-suite/scenarios/run" \
  -H "X-API-Key: $TOKEN" -H "Content-Type: application/json" -d '{
  "platform": "telegram-botapi",
  "scenario": {"name": "order flow", "turns": [
    {"stimulus": {"kind": "command", "text": "/start"},
     "expectation": {"kind": "send_message", "textContains": "Welcome", "buttonContains": "Order"}},
    {"stimulus": {"kind": "callback", "data": "order"},
     "expectation": {"kind": "send_message", "textRegex": "Order created #\\d+"}}
  ]}
}'

Saved scenarios: POST /bot-suite/scenarios, GET /bot-suite/scenarios, POST /bot-suite/scenarios/{id}/run, GET /bot-suite/scenarios/{id}/runs.

Stand endpoints, all namespace-scoped:

Call Does
POST /bot-suite/sessions create a stand — returns sessionId, token, baseUrlAbsolute (optional ttlSeconds)
GET /bot-suite/sessions list live stands with their mode and expiry
GET /bot-suite/sessions/{id} one stand: mode, webhook state, recent deliveries
POST /bot-suite/sessions/{id}/updates send an update — {kind, text, data, chatId}; in webhook mode the answer reports whether it reached the bot
GET /bot-suite/sessions/{id}/actions what the bot did, plus the delivery log
POST /bot-suite/sessions/{id}/webhook register the bot’s webhook ({url, secretToken, allowedUpdates, maxConnections})
DELETE /bot-suite/sessions/{id}/webhook back to long-polling
DELETE /bot-suite/sessions/{id} close the stand
# Create a stand and talk to a bot pointed at it
curl -X POST "http://localhost:5770/api/v1/namespaces/my-namespace/bot-suite/sessions" \
  -H "X-API-Key: $TOKEN" -H "Content-Type: application/json" -d '{"ttlSeconds": 3600}'
# → {"sessionId":"...","token":"...","baseUrlAbsolute":"http://localhost:5770/botapi/..."}

curl -X POST "http://localhost:5770/api/v1/namespaces/my-namespace/bot-suite/sessions/$ID/updates" \
  -H "X-API-Key: $TOKEN" -H "Content-Type: application/json" -d '{"kind": "command", "text": "/start"}'

curl "http://localhost:5770/api/v1/namespaces/my-namespace/bot-suite/sessions/$ID/actions" \
  -H "X-API-Key: $TOKEN"

From an AI agent

The MCP tool bot_dialog_run runs an inline scenario and returns the transcript in one call; bot_scenario_create / bot_scenario_run_saved / bot_dialog_get_run manage saved scenarios.

For the stand: bot_stand_create returns the URL and token to configure the bot with, bot_stand_send_update acts as the user, bot_stand_actions reads what the bot did, bot_stand_set_webhook switches it to webhook delivery, and bot_stand_list / bot_stand_get / bot_stand_close manage the stands. That is a full bot-testing loop without a human.

In a Test Plan

Add a bot_scenario item referencing a saved scenario — the step passes when the dialog passes and fails (with the transcript in the step report) otherwise.

If the scenario expects the bot to answer (any expectation other than silence), the item must name the stand the bot is polling: put {"sessionId": "<stand id>"} in the item’s parameters (the stand id from the Bot Suite or bot_stand_list) and keep the bot’s api_url pointed at that stand. Without it the item is refused with a message that says so, instead of running against a stand nobody is connected to and failing every turn. A scenario made only of silence checks runs without a stand.

As a test case step (the accumulated regression)

A bot has no endpoint, so an ordinary case step cannot express anything about it. A step of type bot runs the conversation instead, through the same dialog engine — which is what makes a conversational product’s test suite actually executable, and what lets a recalled set of cases be re-run unattended.

{
  "executorType": "bot",
  "executorConfig": {
    "platform": "telegram-botapi",
    "sessionId": "<the stand the bot's api_url points at>",
    "scenario": {"name": "order flow", "turns": [
      {"stimulus": {"kind": "command", "text": "/start"},
       "expectation": {"kind": "send_message", "textContains": "Welcome"}}
    ]},
    "successAny": ["order created"]
  }
}
  • sessionId is required for telegram-botapi and generic, and it is not created for you: the bot under test must already be polling that stand (its api_url). A step that names none fails with a named error instead of being run against an empty stand — every turn that expects the bot to act would fail and the report would blame your product for our stand.
  • scenario is the dialog, in exactly the shape bot_dialog_run accepts (see A scenario is a list of turns above). A step may name a saved scenario instead ("scenarioId": "<id from Scenarios>") — its platform comes from the saved row, and the run then appears in that scenario’s own run history.
  • successAny is optional and is the criterion for the whole conversation: the step fails if the last thing the bot said contains none of these. A flow whose every turn passed and which ended somewhere else is exactly what this catches.
  • platform defaults to telegram-botapi. memory runs the scripted double and needs no stand.

The step’s verdict is the dialog’s: passed when the flow passed (and the goal, if declared), failed otherwise, with the per-turn transcript on the step report. Nothing else in the case needs changing — a bot step dispatches like any other automated step.

Limits — what this is, and what it is not

The stand is a mock of the Bot API, not Telegram. Be aware of the boundaries:

  • No real Telegram. Nothing is sent to real users, no phone number or bot registration is involved, and messages the bot “sends” go nowhere but the stand’s action log. Files uploaded to it are kept in memory (a few megabytes, most recent first) and disappear with the stand.
  • Only the update kinds listed above. A user is simulated as a message, a /command, an inline-button press, a link press or an inline query. Group/channel events, payments, polls, edited messages and chat-member updates are not simulated (a bot that calls those methods still gets ok: true and the call is recorded).
  • Webhook delivery is bounded, not persistent. A failed delivery is retried a small number of times with backoff and then given up on, with the reason recorded — the real service retries far longer. getWebhookInfo reports the last failure.
  • A stand lives on the node that created it. With several admin nodes behind a load balancer, point the bot at the address of that node; the stand is not shared across the cluster and does not survive a restart.
  • A stand expires when idle — 30 minutes by default, up to 24 hours.
  • A dialog run is bounded: each turn’s reply window is capped at 5 minutes, the poll interval floor is 100 ms, and a scenario is limited in turn count — a run always terminates.