LLM Jailbreak
An LLM jailbreak bypasses a language model's safety mechanisms, making it provide content it should refuse.
An LLM jailbreak is an attempt to bypass a language model's safety guidelines. The goal is to make the model do or reveal things it should refuse.
How it works
Providers and developers define what a model must not do. This includes harmful content, internal instructions and confidential data. In a jailbreak, a user attempts to bypass these limits using clever phrasing. Typical patterns include role-playing or invented scenarios. Step-by-step conversations can also slowly lead the model away from its rules. Unlike prompt injection, the attack usually comes directly from the user.
Example from practice
A company operates a chatbot for customer service. An attacker engages it in a role-playing game. The chatbot pretends to be a developer debugging an issue. After a few interactions, the chatbot reveals its internal instructions. These contain information about connected systems and customer data it can access. The attacker then tries to obtain specific data from other customers.
Why it matters
A jailbreak itself is often harmless if the model has no access to sensitive data or tools. It becomes critical when the chatbot is connected to customer data or internal systems. What begins as a game can quickly turn into a data leak.
What to look out for
- Assume internal instructions will eventually become public. Do not store secrets in them.
- Limit the data the chatbot can access. Check permissions outside of the model itself.
- Filter the model's output for confidential information before it is displayed.
- Monitor conversations for any suspicious patterns.
- Regularly test your application with targeted jailbreak attempts.
Switzerland and regulation
If a chatbot releases personal data to unauthorised parties, it is a data security breach. This falls under the revised Federal Act on Data Protection (revFADP). If the breach is likely to result in a high risk to the persons concerned, a notification to the FDPIC is required.
Typical mistakes
Companies often rely only on the model provider's built-in safety mechanisms. However, these do not cover the risks of their specific application. Another mistake is using chatbots with broad data access. They often lack separate permission checks.
How we implement it
Our Agents operate within the defined limits of your mandate. They submit actions with significant consequences to our analysts for approval.