Monitor LLM drift

Know when a surface’s model starts answering differently, even while every request still succeeds, and test a candidate model against real captured traffic before it takes any live share. Field-level reference for every setting here is in LLM drift →.

GuideWhat you will achieve
Detect LLM drift and set a baselineAdd the Drift Harness to an LLM Surface, pin a baseline of its normal behaviour, and get an alert when refusals, latency, cost, or other dimensions move past your thresholds.
Replay captured traffic against a candidate modelRun a surface’s captured prompts through another surface or a variant with a different model, and compare the answers with the baseline before any live traffic moves to it.