AI

Model Poisoning

Model poisoning manipulates the training or knowledge data of an AI model to alter its behaviour in a targeted way.

Model poisoning is the manipulation of an AI model through its training data or parameters. The model subsequently behaves incorrectly or according to the attacker's intentions in certain situations.

How it works

An attacker introduces manipulated data into the training process. This can happen through public data sources or contributions to open-source datasets. It can also occur via compromised internal data. One variant involves hidden backdoors. The model operates normally but reacts incorrectly to a specific trigger. Pre-trained models from public sources can also be manipulated. For RAG systems, the risk extends to the documents the model uses at runtime.

A practical example: A company trains a model that detects phishing emails. Part of the training data comes from a public collection. An attacker has marked emails with a certain characteristic as harmless. The model then allows exactly these kinds of emails to pass through. An evaluation with an independent test set reveals the vulnerability.

What to look out for

  • Check the origin of training data and pre-trained models.
  • Obtain models only from trusted sources and verify their integrity.
  • Protect internal training data against unauthorised modifications.
  • Test models with independent test sets before putting them into operation.
  • Monitor the model's behaviour for unexpected changes during operation.

Why it is difficult to detect

A poisoned model behaves correctly in most cases. The manipulation only reveals itself under specific conditions. Standard tests therefore often fail to detect the issue. Targeted tests and good documentation of data sources can help.

How it differs from prompt injection and jailbreaking

Prompt injection and jailbreaks attack a model through its inputs. Model poisoning occurs earlier in the process. It targets the training or the data the model relies upon. The effect is more permanent and harder to reverse.

Switzerland and regulation

There are no specific regulations for model poisoning in Switzerland. For companies with EU links, the EU AI Act has requirements for high-risk systems. These include the quality and integrity of training data.

Typical mistakes

Public models and datasets are often used without prior inspection. A second mistake is a lack of version control. This makes it impossible to determine which data shaped a model.

How we implement it

Our agents run on Claude in a dedicated tenant on AWS in Switzerland. Consequential interventions remain bound to your mandate and to approval by our analysts.

How ANOMAL implements this