Instruction Hierarchy
A model-training technique where LLMs are taught to weight instructions by their source in a defined hierarchy—operator system prompt > authorized user input > retrieved documents > tool outputs—so that lower-priority sources cannot override higher-priority ones. Introduced by OpenAI with the o1 model family and adopted (in different forms) by Anthropic's constitutional classifiers and Google's Gemini safety layer. Instruction hierarchy raises the difficulty of Prompt Injection but does not eliminate it: the model still processes all instructions on one channel and can be bypassed by novel phrasings, encoded payloads, or authority claims.
Frequently asked questions
- Is instruction hierarchy a solved defense?
- No. Every frontier lab that has published on it acknowledges it's a mitigation layer, not a solution. Benchmarks post-2024 consistently show that novel injection payloads defeat instruction hierarchy at meaningful rates. It's one layer of a defense-in-depth stack; the load-bearing security is still architectural (least privilege, egress control, sandboxing, output filtering).