Skip to main content
Experiments are in early access. Ask your PolyAI representative to turn them on for your project.
An experiment runs one of your Branches alongside your Live version and splits real customer calls between them. You compare how the two perform, then keep the one that did better. Use one for any change where you want evidence before every customer gets it, such as a new prompt, a reworked flow, a different routing rule or a model swap. Customer calls split between the control (your current Live version) and the variant (the Branch you are testing). Ending the experiment picks a winner that takes every call. Customer calls split between the control (your current Live version) and the variant (the Branch you are testing). Ending the experiment picks a winner that takes every call.

How it works

An experiment always compares two versions. Variant here only means the Branch under test. It has nothing to do with variants, which set per-site behaviour. You choose what share of calls the variant gets, anywhere from 1% to 99%, and the control gets the rest. The default is an even split. Each call is assigned to one version when it starts and stays on that version until it ends. Every call during the experiment is tagged with the experiment and with the version that handled it. That is what lets you filter your dashboards by experiment and compare the two versions side by side. You can keep updating both versions while the experiment runs. When you have enough data, you end the experiment and pick a winner, which then takes every call.

Before you start

You need:
  • Something published to Live. This becomes the control.
  • A Branch with the change you want to test, synced with Live. See Making a change. A Sub-branch can’t be tested on its own, so merge it into its Branch first.
  • No other experiment running on the project.
  • The Deployment permission. See Access control.
Run simulation tests on the Branch first to check the change works. The experiment then shows how it performs with real customers.
Decide what winning means before you start, for example more calls contained without longer call times. If you pick the metric after seeing the results, it is easy to keep a change that made things worse.

Start an experiment

1

Open the New experiment page

On the Deployments page, choose Create, then Experiment. You can also choose Start Experiment from a Branch’s menu on the Deployments page, or from the Publish menu while you are working in the Branch.
2

Name it

The name defaults to the date and time. Change it to something you will recognize later, such as Refund flow rewrite.
3

Choose the Branch to test

Under Versions, Version A is the control, which is always your Live version. For Version B, select the Branch you want to test as the variant and enter its share of traffic. The control gets the rest.
4

Create

Both versions start taking real customer calls straight away.
Both versions are live. Every caller reaches one or the other and has a real conversation. Only test a Branch you would be comfortable giving to every customer.

While an experiment is running

The experiment appears on the Experiments tab of the Deployments page, marked Running, with its traffic split. If you talk to your Live agent from the chat or call panel in Agent Studio, either version may answer. You can keep improving both versions without stopping the experiment. To rename the experiment or change the split, open its menu on the Experiments tab and choose Update. Changing the split part way through makes the two versions harder to compare, so only do it when you need to. A few things wait until the experiment ends:
  • Publishing or archiving the Branch under test.
  • Starting another experiment.
  • Rolling back Live.

Track performance

Because every call is tagged with the experiment and its version, you can compare the two in your dashboards or by asking Wren.

Dashboards

Once experiments are turned on for your project, every dashboard on the Analytics page has an Experiment filter and an Experiment version filter. See Self-serve dashboards.
1

Choose the experiment

Open a dashboard and select the experiment under Experiment. Most charts split by version, showing Control next to your Branch’s name. A chart that already has its own grouping may show both versions combined.
2

Set the dates

Set the time range to cover the period the experiment ran.
3

Look at one version on its own (optional)

Choose it under Experiment version.
Past experiments stay in the Experiment list after they end, so you can go back to their results at any time.

Wren

Wren knows about your experiments, running and finished. Ask it how the versions compare and it works out each metric for both, side by side.
  • “How is the refund flow experiment doing on containment this week?”
  • “Compare call length between the two versions of last month’s experiment.”
Wren gives you the figures for each version but leaves the decision to you, because a small difference can be down to chance. You can also ask Wren to build a dashboard for an experiment, which opens with that experiment already selected. See Analyze conversations. Wren can start, re-split or end an experiment for you too. It always shows you exactly what it will do and waits for your approval first.

End an experiment

1

Open the Experiments tab

It is on the Deployments page.
2

Choose End

Open the running experiment’s menu and choose End.
3

Pick a winner

Select the version to keep, then choose End experiment. If the version you picked is out of sync with Live, the button reads Sync and end experiment, and it syncs the Branch first.
The winner takes every call from then on. If the winning Branch can’t merge because of conflicts, the experiment keeps running. Sync the Branch with Live, resolve the conflicts, then end the experiment again.

Past experiments

The Experiments tab lists every experiment with its traffic split. Finished experiments also show whether Control won or Variant won. Their data stays in your dashboards, and Wren can still answer questions about them. A/B tests from before experiments were turned on are listed too, but without a result.

Limits

  • One experiment per project at a time.
  • Two versions per experiment, your Live version and one Branch.
  • Sub-branches can’t be tested. Merge them into their Branch first.
  • Agent Studio doesn’t tell you whether a difference between the versions is big enough to trust. Your dashboards and Wren give you the figures, and you decide when there is enough data to call a winner.
  • Charts grouped by deployment, and tables with a row grouping, don’t show while you compare versions. For a grouped table, remove the grouping or pick one version under Experiment version.

Branches

Make the change you want to test.

Deployments

How changes move from a Branch to Live.

Self-serve dashboards

Filter any dashboard by experiment and version.

Analyze conversations

Ask Wren how your two versions compare.

Simulation tests

Check a Branch works before real customers reach it.

Compare changes

See exactly what your Branch changes before you test it.
Last modified on October 9, 2026