Guardrail
Guardrails are technical boundaries that define what an AI agent is permitted to do.
A guardrail is a technical boundary that specifies what an AI system may and may not do. Guardrails prevent harmful outputs, unauthorised actions, and the leakage of confidential data.
How it works
Guardrails operate at multiple stages. Before the model, they check inputs for things like hidden instructions. After the model, they check outputs for personal data or inappropriate content. For agents, they also limit actions. This includes which tools are permitted and with which parameters. They also decide when human approval is necessary. Effective guardrails are anchored in code, not just in the model's instructions.
A practical example
An agent in the SOC suggests isolating a server. A guardrail checks if the server is marked as business-critical. If so, the action is not executed. Instead, it is presented to an analyst for approval. For a standard laptop, the agent can act directly within its mandate.
What to look out for
- Do not rely solely on instructions within the prompt. These can be bypassed.
- Define for each action whether it is automatic, requires approval, or is forbidden.
- Log every guardrail decision in the audit trail.
- Test guardrails regularly with targeted attacks.
Typical mistakes
Guardrails are often set up once and then never reviewed again. New tools or data sources, however, expand the agent's capabilities. A second mistake is granting overly permissive rights for the agent's technical access.
Switzerland and regulation
There are no specific regulations for guardrails in Switzerland. However, the revFADP (revised Federal Act on Data Protection) requires appropriate measures to protect personal data. Guardrails that prevent the leakage of such data help to meet this requirement.
How we implement it
Our agents operate within the boundaries of your mandate. Thought processes and evidence are recorded in the Case.