Replay captured traffic against a candidate model

Run a surface’s captured prompts through another surface or a variant with a different model, and compare the answers with the baseline before any live traffic moves to it.

This guide replays the prompts a surface has captured through a candidate, either another LLM Surface or a variant of the same surface, and compares the candidate’s answers and dimensions with the baseline. For how replay evaluation works, see LLM drift detection →.

Synthetic test prompts rarely match what callers actually send, and a canary split shows the candidate’s answers to real callers. A replay runs recent real prompts through the candidate and shows its answers on the LLM Drift page instead of returning them to callers, so you can judge the change before any caller sees it.

Use this guide when:

  • You are about to change a surface’s model, provider, prompt, or Judge and Jury setup, and want evidence from recent real traffic first.
  • You have created a variant for a canary split and want to test it before it takes any share of live traffic.
  • You want to compare how two surfaces answer the same captured prompts.

Prerequisites

  • Drift monitoring on the source surface, with a pinned baseline and some captured traffic. See Detect LLM drift and set a baseline. Prompts are captured only while drift monitoring is on.
  • A candidate to replay through, either:
    • A variant of the source surface: an enabled variant with the candidate model, created as in the first three steps of Roll out a model with a canary split. The source surface’s baseline is the one its results are compared with.
    • Another LLM Surface: a surface with the candidate model and its own Drift Harness with drift monitoring on. Pin a baseline on it too, since its results are compared with its own baseline.
  • An embeddings model and its API key secret available under Models, to measure how far each replayed answer moves from the captured one.
  • The Administrator role.

Each replayed prompt is a real call to the candidate’s provider. It is billed to the target surface and counted against that surface’s usage limits.

Steps

Turn on replay evaluation on the source surface

Under SURFACES in the dashboard sidebar, select LLM and open the source surface. Select Drift Harness on the canvas, then select Configure drift harness…. In the Drift Harness — behavioural drift monitoring card, turn on Replay evaluation.

Choose the embeddings model

Under Embeddings model, select the Embeddings provider, then the Embeddings model, then the API key secret that holds the provider’s key.

Set the run limits

Set Max sessions per run, from 1 to 200, to the number of captured prompts one run replays. It starts at 50. Optionally, enter a Cost ceiling per run in USD to stop a run part-way once it has spent that much. Leave it blank to use the appliance default, or enter 0 for no ceiling.

Save the surface

Select the save button (Save changes) under Manage.

Open the Evaluation tab

Under Monitoring in the dashboard sidebar, select LLM Drift, then select the Evaluation tab. The Indicative surface data card shows how many prompt and response records the selected surface has captured.

Choose the source and the candidate

Set the three selectors:

  • Surface Source Data to Replay: the surface whose captured prompts you want to replay.
  • Evaluation Target Surface: the candidate surface, or the source surface itself when the candidate is one of its variants.
  • Target Variant: the candidate variant, or Base variant when the candidate is another surface.

The line below the selectors confirms the run, for example “Replays source’s 50 records through target — an experiment, so source’s baseline stays unchanged.”

Run the experiment

Select Run experiment. Progress shows as “Replaying done/total”, with a Cancel button to stop the run.

Read the result

When the run completes, the Evaluation tab shows:

  • A Stable percentage: the share of replayed answers that stayed close to the captured answer.
  • The Semantic distance: how far, on average, the replayed answers moved from the captured ones.
  • A table of each dimension with its Baseline, This run, Change, and Status, judged against the target surface’s baseline.
  • One chart per dimension, with this run’s value in amber against the baseline’s dashed line.

A Drift status means this run’s value for that dimension moved past the target surface’s Drift threshold, relative to its baseline.

Confirm

Test 1: a manual experiment run returns a History row with an Experiment badge

Select the History tab. Expected: the run is listed under Evaluation history with an Experiment badge, a Manual trigger, and a Completed status. Select the row to open its detail panel, which shows Source data from Surface with the source surface’s name.

Test 2: a self-check run returns the source surface’s own figures for comparison

On the Evaluation tab, select the arrow button between Surface Source Data to Replay and Evaluation Target Surface to point the target back at the source, then select Run evaluation. Expected: a second result for the source surface itself. Compare its Stable percentage and Semantic distance with the experiment’s. A self-check run also adds its figures to the source surface’s charts on the Drift over time tab, while an experiment leaves those charts unchanged.

Troubleshooting

SymptomLikely causeFix
Run experiment or Run evaluation is greyed out, and the tab reads “No records captured yet.”The source surface has captured no prompts, because drift monitoring was off or no traffic has arrived.Turn on Monitor drift on the source surface’s Drift Harness, save, and send traffic to it.
The candidate surface is missing from Evaluation Target Surface.Only surfaces with drift monitoring on are listed.Add a Drift Harness to the candidate surface and save it.
Target Variant is greyed out and shows only Base variant.The target surface has no enabled variants.Turn on the variant’s Enabled switch in the surface’s Variants tab and save, or create a variant.
Semantic distance reads n/a.No embeddings model is set on the source surface’s Drift Harness.Choose an Embeddings provider, Embeddings model, and API key secret under Replay evaluation, and save the source surface.
The run ends with Stopped on cost.The run reached the Cost ceiling per run, or the target surface reached a usage limit.Raise the Cost ceiling per run or lower Max sessions per run, then run again.
The run ends with Failed. and “all N replayed sessions failed”.Every replayed prompt returned an error from the target, for example a missing provider key or an unavailable model.Send a request to the target surface or variant directly, fix the error it returns, then run again.
Every dimension’s Status reads No baseline.The target surface has no pinned baseline.Pin a baseline on the target surface, as in Detect LLM drift and set a baseline.

Next steps