The Results tab of the Conversation Simulation page lists every simulation run and shows which agent answers matched the expected responses. Found in the agent sidebar under Simulation β Scenarios, then the Results tab. Opening a run shows its Simulation Result page.
Results Tab: Controls
| |
|---|
| Show Filters / Hide Filters | Opens or closes the filter panel. |
| Sort By | Orders the list by Run Date (default, newest first), Code, Status or Duration, Ascending or Descending. |
| Grid View / List View | Shows the runs as cards or as a table (columns Run, Status, Started, Duration, Conversations, Steps). |
Results Tab: Filters
| |
|---|
| Code | Shows only runs whose code matches what you type. |
| Status | Shows only runs with the selected status: Completed, Failed, Running or Stopped. |
| Run Date From / Run Date To | Shows only runs started within this date range. Both dates are inclusive (whole days). |
| Apply | Applies the filter values to the list. |
| Reset | Clears the filters and shows all runs again. |
Results Tab: Run List
| |
|---|
| Run card or row | One per run. Click it to open the Simulation Result page. |
| Code | The run's identifier, shown as #<code>. |
| Status | Completed, Failed, Running, Stopped or Not Yet Run. |
| Run date and duration | When the run started and how long it took (- while not available). |
| Conversations and Steps | How many conversation simulations and how many steps the run contains. |
| Load More | Shown when more runs are available. Adds the next batch to the list. |
| Empty state | No simulation results yet when nothing has been run, or No Results Found when the filters match nothing. |
| |
|---|
| Back | Returns to the Results tab. |
| Title | Simulation Result with the run code and start time. A small spinner shows while the page is refreshing. |
| Error | Shown in red when the run failed, with the reason and, where relevant, "Running it again may help." |
| Refresh | Shown while the run is Running. Reloads the latest progress. The page also refreshes by itself while the run is in progress; after 10 minutes live updates pause and you use Refresh instead. |
| Stop | Shown while the run is Running. Stops the run after you confirm. |
| Run Again | Shown when the run is not running. Starts a new run of the same conversation simulations and opens it. |
| Delete (trash icon) | Shown when the run is not running. Deletes this run after you confirm. This cannot be undone. |
Simulation Result: Summary Cards
| |
|---|
| STATUS | The run's current status. |
| CRITERIA MATCH | How many evaluation criteria were matched out of the total across all steps, for example "8/10 Matched". |
| DURATION | Total run time. |
| STEPS | Total number of steps in the run. |
Simulation Result: Conversations
One expandable row per conversation simulation in the run.
| |
|---|
| Conversation row | Shows the status icon, conversation name, folder, how many steps were completed (for example "3/4 Steps completed") and its duration. |
| Failed at step #n | Shown when the conversation failed, pointing at the step where it stopped. The error reason is shown below the row. |
| n blocked calls | Shown when tool calls in this conversation were blocked. |
| Mocks used in this run | Opens a panel listing the mock sets the run used and the tools that were allowed to run live. Only shown when the Mocks feature is enabled. |
| Match badge | Shown for completed conversations: criteria matched out of the total for this conversation. |
| Empty state | No Conversations Found when the run has no conversation data. |
Simulation Result: Step Details
Expand a conversation to see each step as its own expandable card.
| |
|---|
| Step #n | The step number with its status icon and, once completed, a match badge for that step. |
| Error | Shown in red when the step failed, with the error message. |
| User message | The message sent to the agent, shown at the top of the card. |
| Expected Response | The answer you wrote in the scenario. |
| AI Agent Response | The agent's actual answer. Parts that support a criterion are highlighted in green, parts that contradict one in red. Knowledge references in the answer can be expanded to show the source text. |
| Evaluation Criteria (x/y Matched) | Shown for completed steps; expand or collapse it. Each criterion is listed with either the quoted excerpt that matched it (green) or the reason it was not matched. |
| No Evaluation Criteria | Shown when the step had no criteria defined. |
Shown inside each step card. Says No tool calls in this turn. when the agent called no tools.
| |
|---|
| Tool calls (n) | Expands the list of tool calls the agent made in this step, with a count per outcome. |
| Tool call row | Shows the call number, the tool, an outcome chip (Mocked, Default, Simulated, Live, Blocked, Internal or Unknown), where the answer came from (for example Step or Conversation mock set), a short preview of the arguments, and a one-line explanation. Live calls carry a warning chip that they may contain real data. |
| Expand a row | Shows the full Arguments and Response of the call. |
| Mock this call | On a Blocked call. Lets you pick where the new mock goes (same choices as Save as case), then opens the conversation form with the mock set editor. Only shown when the Mocks feature is enabled. |
| Save as case | On a Mocked or Live call. Opens the conversation form to save the call as a mock case in a new private set for this step only, a new private set for the whole conversation, or an Existing set⦠attached to the conversation. Only shown when the Mocks feature is enabled. |
Related guide: Run and Track