Guardrails
Rules, filters, and constraints that keep AI agents within safe operating boundaries. Guardrails prevent agents from hallucinating, leaking sensitive data, or taking unauthorized actions. Examples include topic restrictions, PII redaction, confidence thresholds, Stopping Conditions, and Human-in-the-Loop (HITL) approval gates.
Example
A healthcare AI agent has guardrails that prevent it from providing medical diagnoses, redirect clinical questions to a physician, and redact any patient identifiers from responses.
Frequently asked questions
- What types of guardrails exist for AI agents?
- Common guardrails include: topic restrictions (only discuss X), PII redaction (never reveal personal data), confidence thresholds (escalate when uncertain), action limits (require approval for irreversible actions), and content filters (block harmful output).