Skip to main content

Guardrails

AI Guardrails

Pronunciation
GARD-railz
Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/guardrails

In short

Guardrails are checks placed around an AI model that screen what goes in and what comes out, blocking unsafe, off-topic or malformed requests and answers.

What are AI guardrails?

Guardrails are the rules and checks an application wraps around a language model to keep it within bounds. Input guardrails look at a request before it reaches the model, for example to refuse harmful or off-topic questions, strip out personal data or flag a likely prompt injection. Output guardrails inspect the answer before anyone sees it or acts on it: they can block toxic text, catch leaked secrets, check that the answer is backed by the retrieved sources, or make sure the output is valid JSON in the expected format.

They are usually built in layers, cheapest first. Plain code handles exact rules, such as allow-lists, a regular expression for card numbers or a JSON schema check. Safety classifiers, such as Meta's Llama Guard, label text as safe or unsafe across categories of harm, and a second LLM can judge subtler things, such as tone or whether an answer stays on topic. Open-source toolkits such as NVIDIA NeMo Guardrails, and the AI safety services of the large cloud providers, package these checks.

For AI agents, guardrails also limit actions: which tools the model may call, how much it may spend and which steps need a person's approval. Like the guardrails on a mountain road, they don't steer; they stop the worst outcomes when something goes wrong. Each check adds latency and cost, and overly strict ones block legitimate users, so teams tune them against real traffic and their evals.

Guardrails are often confused with alignment and with the system prompt. Alignment training, such as RLHF, changes the model itself so that it tends to behave well; a system prompt asks the model to follow rules it may still ignore. Guardrails sit outside the model and are enforced by the application's code, which makes them a key layer against prompt injection and hallucination, but not a complete fix for either, since attackers keep finding phrasings that filters miss.

Key takeaways

  • Guardrails are checks around an AI model, not part of the model itself.
  • Input guardrails screen requests; output guardrails screen answers before they are shown or used.
  • They combine code rules, safety classifiers and LLM judges, cheapest first.
  • For agents they also limit tools, spending and actions that need human approval.
  • They reduce prompt injection and hallucination risks but can't remove them.

Example

Input and output guardrails around an LLM callpython
def answer(question: str) -> str:
    # Input guardrail: a small, fast classifier screens the request first
    if safety_classifier.is_unsafe(question) or is_off_topic(question):
        return "Sorry, I can only help with questions about your account."

    reply = llm.generate(system=SYSTEM_PROMPT, user=question)

    # Output guardrails: enforced in code, whatever the model was told
    if contains_personal_data(reply) or not supported_by_sources(reply):
        return "I can't answer that reliably. A support agent will follow up."
    return reply

Readers ask

What is the difference between guardrails and a system prompt?

A system prompt is an instruction to the model, which it usually follows but can be talked out of. Guardrails are checks in the application's code that run whatever the model does, so they still apply when the instructions fail.

Do guardrails stop prompt injection?

They help but don't stop it completely. Input filters catch many known attack patterns, and output and action checks limit the damage when one gets through, which is why they are combined with least privilege and human confirmation for sensitive actions.

Do guardrails slow down an AI app?

Each check adds some delay and cost, especially when it calls another model. Teams run fast code checks first, run model-based checks in parallel with the main call where they can, and save the heaviest checks for risky requests.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

More

Settings