Model Cascade
A cost-optimization pattern where a cheaper, faster model handles the first attempt at a task, and a more capable (but slower and more expensive) model is invoked only when the first model fails a confidence threshold or quality check. Cascades typically cut total cost 60-90% on workloads dominated by easy cases (a large fraction of support tickets, document classification, simple coding tasks), because most queries never reach the expensive model.
Example
A customer-service AI agent fronts every query with a small model (Haiku, GPT-5-mini) that handles 70% of requests directly. For the 30% where the small model is uncertain—measured by output confidence, retrieval scores, or a self-check—the request escalates to a frontier model (Claude Opus 4, GPT-5). Average per-ticket cost drops from $0.18 to $0.06 with no measurable quality loss on the cascade-handled cases.
Frequently asked questions
- When does a model cascade NOT pay off?
- When the task distribution is mostly hard cases (the cheap model fails on 60%+ of inputs), when latency budget is tight (cascading adds a second model call when escalating), or when the cheap model produces *plausibly wrong* answers the system can't easily detect as wrong. The cascade only wins when easy cases dominate AND the cheap model's failures are reliably caught.
- How do you decide when to escalate?
- Common signals: low retrieval scores, the cheap model's own self-rated confidence, mismatched outputs across two passes of the cheap model, or a fast LLM-as-judge check on the cheap model's output. Threshold tuning is empirical—start with a conservative threshold (escalate often), measure quality and cost, then tighten until you hit the cost-quality target.