Reliability without effort
A single provider outage, rate limit, or degraded model should not take a feature offline. Yet building retry, failover, and caching into every agent and service is repetitive work, and each team ends up reimplementing it inconsistently, or not at all.
What you can do
- Fail over to backup providers of a different vendor or model so that a single provider outage degrades performance instead of breaking the feature.
- Load-balance across a pool of providers by lowest latency or lowest cost so that traffic follows the fastest or cheapest healthy option.
- Retry transient failures with exponential backoff and trip a circuit breaker on a degraded provider so that one bad upstream does not cascade.
- Serve repeat and near-duplicate prompts from an in-memory cache so that common requests skip the provider call entirely and return faster.
- Apply all of this per surface with no external infrastructure so that no team re-implements resilience in its own code.
How it works
Before a request reaches a provider, the surface can serve it from an in-memory response cache. An Exact Cache matches a hash of the normalised request, and a Semantic Cache matches paraphrased duplicates above a similarity threshold, so a cache hit returns without any provider call. On a miss, the call is protected by a Retry Policy with exponential backoff and a Circuit Breaker that opens on repeated upstream failures. Failover then retries against ordered backup providers of a different vendor or model, stopping at the first that succeeds, while a Load Balancing Pool can instead distribute traffic by weighted split, lowest observed latency, or lowest catalogue cost. Failover and load balancing are configured per surface and are mutually exclusive at runtime.
Failover is content-safety-aware: a request blocked by a guardrail is never silently retried against another vendor, and if every eligible backup fails, the request fails rather than returning a degraded or unsafe answer.
Related
- Resilience and caching: How failover, load balancing, circuit breakers, and the response cache work together on a surface.
- Fail over to a backup provider: Step-by-step configuration of an ordered failover chain across vendors.
- Balance traffic across providers: How to distribute requests across a pool by latency or cost.
- Cache repeated responses: How to turn on exact and semantic caching and measure the spend it avoids.
Glad to hear it! Please tell us how we can improve more.
Sorry to hear that. Please tell us how we can improve.
Thank you for sharing your feedback so we can improve your experience.