Test Tools Safely with Mocks

Open in CMS

Once your agent calls tools (your own API, an MCP server, another agent or a plugin), testing gets harder: a test run could book a real table, read live data that changes every day, or fail because a service is down. Mocks fix that. In this step you write scripted answers for your tools, attach them to a conversation simulation, and run it, so every run gets the same predictable tool results and nothing real is touched.


Before you start


Why use mocks

During a simulation, every tool call goes through a gate that decides what happens to it. Without mocks, each kind of tool has a fixed default:

ToolDefault in a simulation (no mock)
Custom API, MCP server, other agents and A2A peersBlocked: no call is made and the agent is told the tool is unavailable in this run.
Plugins that read data (for example Dynamic Table or Public Documents)Live: the real call runs and returns real data.
Plugins and built-in tools that write data (forms, reminders, escalation)Simulated: no real call; the agent gets an empty "success" answer.
Web search and web scrapingLive, and they cannot be mocked.

So without mocks, a scenario that depends on your booking API can never pass, and a scenario that reads a live table can pass today and fail tomorrow. A mock gives the tool a scripted answer instead. That lets you:

  • Test without touching real systems or real data. No real bookings, no real customer records, no calls to a paid API.
  • Get the same result on every run, so a failing run means the agent changed, not the data.
  • Test the hard paths on purpose: an empty result, a "not found", a 500 error, a timeout, a sold-out item. These are difficult to trigger with a real service.

Key terms

TermMeaning
Mock setA named collection of scripted answers for one or more tools. Library sets are shared and can be reused by any conversation simulation of the agent.
Tool mockThe part of a mock set for one tool.
CaseOne scripted answer for a tool, used when the call's arguments match the case's conditions. A tool can have up to 50 cases, checked from top to bottom.
ConditionA rule on one argument of the call, such as query contains latte. All conditions of a case must hold.
OtherwiseWhat a tool answers when no case matches: a default response, or "keep looking" in the next mock set.

Step 1: Create a mock set

  1. In the CMS, open Simulation → Mocks in the agent sidebar. The page is titled Tool Mocks.
  2. Click New Mock Set.
  3. Enter a Name (required, up to 120 characters), for example "Menu: Iced Caramel Latte sold out", and an optional Description.
Mock set editor with the name, description, the Tools in this set list and the selected tool's details

The Library chip next to the title means this set is shared: you can attach it to any conversation simulation of this agent.


Step 2: Add the tool you want to mock

  1. Under Tools in this set, click Add tool.
  2. Search for the tool and click it. The list shows the tools of the published agent, grouped by source (Custom API, MCP, Plugins, Built-in), with what each tool would do if you don't mock it, such as Simulated success.
Add tool list filtered to escalate_conversation, showing the Writes data and Simulated success labels

The selected tool's card shows its description and its Inputs (the arguments the AI fills in when it calls the tool). You can add several tools to one set; select a tool in the list on the left to edit its cases.

If a tool is later removed from the published agent, the set shows Tool not found in the published agent. Use Re-bind to… to move its cases to another tool.


Step 3: Add a case with conditions

  1. Click Add case. Give it a label so you recognise it later, for example "Latte is sold out".
  2. Under When all of these match, click Add condition, choose the argument, choose a comparison and type the value.
  3. Add more conditions if needed. The case only answers when every condition holds. Arguments you don't list can have any value.
A case labelled Latte is sold out with the condition query contains latte and a table response with one row

The comparisons you can choose depend on the argument's type:

Argument typeComparisons
Text (also lists, objects and unknown types)is, is not, contains, starts with, ends with, matches pattern, is one of, is empty, is not empty
Number=, ≠, <, ≤, >, ≥, is between (inclusive)
Yes/nois true, is false
Any typeis present, is missing

Tips:

  • Text comparisons ignore upper and lower case unless you switch on Aa (match case) on the condition.
  • is one of takes a comma-separated list; is between takes a minimum and a maximum.
  • matches pattern uses a regular expression.
  • A case with no conditions matches every call. Put it last, because cases are checked from top to bottom and a catch-all would hide the cases below it.
  • For unusual inputs, open Advanced — Match raw arguments and paste a piece of JSON the call's arguments must contain. It is checked in addition to the conditions.

Step 4: Write the response

Under Respond with, write what the tool returns for this case. The editor adapts to the tool:

  • Form shows a form that fits the tool's output (for example rows and columns for a table query); View as JSON shows the raw body. You can also switch between JSON and Text.
  • Success and Error open ready-made templates, such as an empty list, "Not found (404)", "Unauthorized (401)", "Rate limited (429)", "Service down or timed out" for a Custom API, or "Query fails" and "No rows found" for a table query.
  • For a Custom API, choose the HTTP Status (default 200 OK). Error codes in the 5xx range reach the agent as an empty result, just like a real failing API.
  • For an MCP tool, choose Content or MCP error. For another agent, write the reply text.
  • Suggest lets AI draft a response for the case. Format tidies the JSON.
  • Fetch a sample (tools that only read data) calls the real tool once so you can copy its shape. Use a real response is only available for Custom APIs that use GET; for other tools it is locked because calling them could create real data.

Response bodies can be up to 64 KB.


Step 5: Decide what happens to every other call

Below the cases, the Otherwise (every other call) box decides what the tool answers when no case matches. Open Advanced to choose:

  • Respond with a default: the call gets the response you write in this box (here: "No rows found").
  • Keep looking: the call moves on to the next mock set, and if none answers, to the tool's default from the table in Why use mocks. In the tool list such tools are marked "Keeps looking".
Otherwise box with an empty table response and the Advanced options Respond with a default and Keep looking

Click Save. The set now appears on the Tool Mocks page with its tools, number of cases and how many conversations use it.

Tool Mocks page with one mock set card showing 1 tool, 1 case and Used in 1 conversation

Step 6: Attach the mock set to a conversation simulation

  1. Open Simulation → Scenarios and open (or create) a conversation simulation. See Design Scenarios for writing steps and criteria.
  2. In the Tool mocks card, click Add mock set:
    • From library… attaches a shared set, such as the one you just made.
    • Create new makes a Private set that belongs only to this conversation.
  3. Click Save on the conversation.
Tool mocks card on a conversation simulation with one Library mock set attached and the Date and time the agent sees switch

Things to know:

  • Sets attached here are used by every step. Drag the rows to change their order: the first set that answers a call wins.
  • Each step also has its own Tool mocks (n) row with Add step mock set. A step's own sets are checked before the conversation's sets, so you can change one tool's answer for a single step (for example "the first search finds nothing, the second one finds the item").
  • Date & time the agent sees: turn the switch on to fix the date, time and time zone for this simulation. Use it when tool arguments depend on dates ("tomorrow at 7"), so the conditions keep matching on every run.
  • In the row's ⋮ menu, Make a private copy (on a library set) gives this conversation its own copy that you can change without affecting other conversations, and Save to library (on a private set) shares it with the rest of the agent's simulations.
  • When you click Run and the run would hit a blocked tool, the confirmation lists it with a Mock it link, which opens the mock editor for that tool.

Check the result

  1. Click Run on the conversation and confirm. When the run is done, open it from the Results tab (see Run and Track).
  2. Expand the conversation. Mocks used in this run opens a panel with the exact mock sets the run used, frozen at the time of the run, and the simulated clock if you set one.
Conversation row in a simulation result with the Mocks used in this run link and 2/2 Matched
  1. Expand a step and open Tool calls. Each call shows an outcome chip and where its answer came from:
Tool call row with the Mocked and Conversation chips, the call arguments and the matched case and mock set
OutcomeWhat happened
MockedA case matched. The line underneath names the case and the mock set, plus the status code. The second chip says whether the set came from the Step or the Conversation.
DefaultNo case matched, and the tool's Otherwise default answered.
SimulatedNo mock answered; the tool's built-in simulated success was used.
LiveThe real tool ran. These rows carry a warning that they may contain real data.
BlockedNo call was made. Use Mock this call to add a mock for it.
InternalAn internal tool that always runs.

Expand a call to see the full Arguments and Response. If the tool has a Response Prompt, the row notes that it was applied on top of the mocked response, just as in production. On a Mocked or Live call, Save as case turns the call into a new case (in a private set for this step, for the whole conversation, or in an existing set), which is the quickest way to build mocks from real behaviour.

In the Brew demo above, the user asks "Can I order an iced caramel latte?", the agent searches the menu with query: "iced caramel latte", the "Latte is sold out" case answers, and the agent replies that the drink is unavailable. Every run gets the same answer, whatever the live menu says.


Limits and good practice

  • Mocks only apply to conversation simulations. Preview, the published agent and real channels always call the real tools.
  • A mock set holds up to 50 tools, each tool up to 50 cases, each case up to 20 conditions.
  • Web search and web scraping cannot be mocked; they run live in simulations.
  • Mocks follow the published agent: publish after adding or renaming a tool, then update your mock sets.
  • Keep one purpose per mock set ("happy path", "API down", "item sold out") and combine sets per scenario, rather than one huge set.
  • Always test at least one error path (Error templates) so you know how your agent answers when a service fails.

For the full list of fields on these screens, see the Help pages Mocks, Scenarios and Results.


Next steps