Shut down the OWASP top risks
Every AI feature you add widens the attack surface. Prompt injection sits at the top of the OWASP list for LLM applications, yet a single content filter is a single point of failure: miss one jailbreak, or let one classifier go offline, and nothing stands between a crafted request and the model. As more teams ship their own features, each wiring up its own checks, safety becomes inconsistent and impossible to prove.
What you can do
- Apply one auditable safety configuration across every model and team so that protection does not depend on each team getting its own checks right.
- Block prompt injection and jailbreak attempts before the model runs so that an unsafe request never reaches a provider or incurs a model cost.
- Mask or pseudonymise PII and secrets on the way in so that sensitive data is never handed to a provider in the clear.
- Fail closed by default so that a classifier outage blocks the request rather than silently passing it through unprotected.
- Label every block with a rejection category so that you can show exactly what was caught, and why.
How it works
Safety is configured once on a surface and applied to every call through a layered pipeline, so no single control is the only thing standing between a request and the model. Prompt Guard runs deterministic regex and model-backed inspection on both seams, rejecting a match outright or applying redact-and-continue to replace it with a static placeholder or a stable NER ID. Expert Witnesses add purpose-trained external classifiers on the request seam, the response seam, or both. On the request seam the Judge runs an LLM-based pre-check in parallel with the Decider, blocking a clearly unsafe request before the main model is called. On the response seam the Jury runs a multi-juror review that can approve, block, or re-drive the answer before it returns.
Every guardrail fails closed by default: if a classifier cannot be reached or a rule engine cannot complete, the request is blocked rather than passed through, and each block is tagged with a canonical rejection category (prompt injection, jailbreak, PII, secrets, toxicity, policy violation, and more) that feeds the rejection dashboards.
Related
- Guardrails: How Prompt Guard, Expert Witnesses, Judge, and Jury layer around a call and why they fail closed.
- PII protection: How sensitive data is detected and pseudonymised with NER IDs before it reaches a provider.
- Block and re-review calls with Judge and Jury: Step-by-step configuration of the request pre-check and response review on a surface.
Glad to hear it! Please tell us how we can improve more.
Sorry to hear that. Please tell us how we can improve.
Thank you for sharing your feedback so we can improve your experience.