Guardrails
AI Guardrails
- Pronunciation
- GARD-railz
In short
Guardrails are checks placed around an AI model that screen what goes in and what comes out, blocking unsafe, off-topic or malformed requests and answers.
What are AI guardrails?
Guardrails are the rules and checks an application wraps around a language model to keep it within bounds. Input guardrails look at a request before it reaches the model, for example to refuse harmful or off-topic questions, strip out personal data or flag a likely prompt injection. Output guardrails inspect the answer before anyone sees it or acts on it: they can block toxic text, catch leaked secrets, check that the answer is backed by the retrieved sources, or make sure the output is valid JSON in the expected format.
They are usually built in layers, cheapest first. Plain code handles exact rules, such as allow-lists, a regular expression for card numbers or a JSON schema check. Safety classifiers, such as Meta's Llama Guard, label text as safe or unsafe across categories of harm, and a second LLM can judge subtler things, such as tone or whether an answer stays on topic. Open-source toolkits such as NVIDIA NeMo Guardrails, and the AI safety services of the large cloud providers, package these checks.
For AI agents, guardrails also limit actions: which tools the model may call, how much it may spend and which steps need a person's approval. Like the guardrails on a mountain road, they don't steer; they stop the worst outcomes when something goes wrong. Each check adds latency and cost, and overly strict ones block legitimate users, so teams tune them against real traffic and their evals.
Guardrails are often confused with alignment and with the system prompt. Alignment training, such as RLHF, changes the model itself so that it tends to behave well; a system prompt asks the model to follow rules it may still ignore. Guardrails sit outside the model and are enforced by the application's code, which makes them a key layer against prompt injection and hallucination, but not a complete fix for either, since attackers keep finding phrasings that filters miss.
Key takeaways
- Guardrails are checks around an AI model, not part of the model itself.
- Input guardrails screen requests; output guardrails screen answers before they are shown or used.
- They combine code rules, safety classifiers and LLM judges, cheapest first.
- For agents they also limit tools, spending and actions that need human approval.
- They reduce prompt injection and hallucination risks but can't remove them.
Example
def answer(question: str) -> str:
# Input guardrail: a small, fast classifier screens the request first
if safety_classifier.is_unsafe(question) or is_off_topic(question):
return "Sorry, I can only help with questions about your account."
reply = llm.generate(system=SYSTEM_PROMPT, user=question)
# Output guardrails: enforced in code, whatever the model was told
if contains_personal_data(reply) or not supported_by_sources(reply):
return "I can't answer that reliably. A support agent will follow up."
return replyReaders ask
What is the difference between guardrails and a system prompt?
A system prompt is an instruction to the model, which it usually follows but can be talked out of. Guardrails are checks in the application's code that run whatever the model does, so they still apply when the instructions fail.
Do guardrails stop prompt injection?
They help but don't stop it completely. Input filters catch many known attack patterns, and output and action checks limit the damage when one gets through, which is why they are combined with least privilege and human confirmation for sensitive actions.
Do guardrails slow down an AI app?
Each check adds some delay and cost, especially when it calls another model. Teams run fast code checks first, run model-based checks in parallel with the main call where they can, and save the heaviest checks for risky requests.
See also
- Prompt InjectionSecurity, p. 35Prompt injection is an attack on LLM apps where attacker-written text is treated as instructions, so the model ignores its rules, leaks data or misuses tools.
- HallucinationAI & Machine Learning, p. 25A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- AI AlignmentAI & Machine Learning, p. 3AI alignment is the field of making AI systems pursue the goals and values their designers intend, so they behave helpfully, honestly, and safely.
- System PromptAI & Machine Learning, p. 48A system prompt is the instructions an app gives a language model before the conversation starts, setting its role, rules, tone and what it should know.
- AI AgentAI & Machine Learning, p. 2An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
- LLMAI & Machine Learning, p. 27An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
Sources
Spotted a mistake or something missing on this page?Suggest an edit