Back in May, the National Cyber Security Centre (NCSC) issued its second assessment of the AI threat to UK organizations. Over the next two years, the technology will “almost certainly … lead to an increase in frequency and intensity of cyber threats”, it warned.

Most news sites leapt on the report’s predictions that AI will help threat actors supercharge phishing, vulnerability exploitation and other offensive activities. But hidden away is a warning of equal importance.

The NCSC predicts that the “growing incorporation of AI models and systems across the UK’s technology base”, especially Critical Infrastructure (CNI) providers, will “almost certainly” increase their attack surface. Here, it is data poisoning that arguably poses the greatest threat.
 
The Threat Is Here

According to the OWASP LLM Top 10, data poisoning occurs when threat actors manipulate “pre-training, fine-tuning or embedding data” in order to introduce vulnerabilities, backdoors and biases. Their goal could be to get the large language model (LLM) to spout inaccurate, misleading or offensive information. Or to degrade its performance. Or to exploit downstream systems that rely on the model’s output, like coding assistants, security tools and autonomous vehicles.

According to the NCSC’s assessment: “AI technology is increasingly connected to company systems, data, and operational technology for tasks. Threat actors will almost certainly exploit this additional threat vector.” However, this is no theoretical threat. It’s already with us. A recent piece of IO research is compiled  from data from 3,000 UK and American information security leaders. We found that 26% had suffered a data poisoning incident over the past 12 months. Separately, IBM reports that 15% of security incidents involving AI models over the past year are down to data poisoning.  

Direct & Indirect Attacks

OWASP claims the risk of data poisoning is particularly high regarding external data sources that may contain “unverified or malicious content”. This is certainly true, but that’s not to say it’s the only risk to consider. In fact, data poisoning takes many forms.

It could be an indirect attack targeting public data on which an LLM is trained on or uses during inference. This is generally the main threat to LLMs of the sort ChatGPT uses. The goal is usually to either degrade the model or introduce a “trigger” which causes the model to ignore its safety guardrails. Indirect attacks might also seek to poison large volumes of data publicly available on the web, in order to make a model less accurate or reliable.
In these attacks, threat actors place malicious data/content in websites or online documents and hope it’s swept up by a model in time. However, data poisoning can also be more direct. This is the type of attack typically targeted at closed LLMs. The end goal may be to extract sensitive corporate data, sabotage the model, force it to spout misinformation, or possibly manipulate it with secret triggers in order to bypass security filters or create malicious code.

How Data Poisoning Works

How do threat actors achieve this? If they’re targeting a data training pipeline directly, they’ll need access. That might come via a malicious insider, or a straightforward network intrusion – via vulnerability exploitation, use of compromised credentials, or other means. They might instead look for the path of least resistance – third-party data sources used to train the model that are managed by an insecure supply chain partner.

With access, they are then able to set about their work. Adversaries might seek to alter existing training data, inject malicious/biased/misleading content into a model’s training data set, or even delete it. They might also mislabel it in order to confuse the model.  

Spotting The Warning Signs

The challenge facing organizations running LLMs is that it can be tricky understanding when a dataset has been compromised. As models are evolving all the time anyway, subtle changes are likely to fall under the radar, especially if the threat actor managed to gain access without setting off any internal alarms.

However, there are some red flags which IT and security teams should take notice of. Unexpected behaviour and unintended outputs should be investigated, as should any increase in false positives or negatives.

If the model produces biased results, that may also indicate a data poisoning incident, as might general performance degradation over time. It goes without saying that corporate security breaches should also be thoroughly investigated for signs that the perpetrators may have been targeting access to LLM training pipelines. The same goes for suspicious employee behaviour.  

Security Versus Innovation

The NCSC warns in its AI threat assessment that the race to gain competitive advantage from LLMs may introduce risks of its own, if it means developers cut corners. With that in mind, it’s critical that LLMs, and the infrastructure that supports them, are built from the ground up with security in mind. This means strong encryption for data at rest and in transit, and best practice identity and access management (IAM), including multi-factor authentication and least privilege policies.

Training or fine-tuning environments could also be isolated to further reduce the risk of unauthorized access.

And supply chain providers must be vetted and audited to ensure strong security posture. The data itself should be sanitised before use to weed out anything that might be malicious, and data streams and model output could be monitored for suspicious activity. Adversarial testing of models will flag any weaknesses that must be addressed to improve resilience.
 
Security and business leaders should educate teams about the risks of AI data poisoning but should also put structured governance and frameworks in place such as ISO 27001 (for information security management) and ISO 42001 (for AI management). They offer a systematic approach to identify security gaps and potential threats and then deliver a framework to address them.

ISO 27001 is foundational for good information security, helping address areas that may affect your AI attack surface, like access controls, data protection and supplier security. And ISO 42001 has been designed specifically to help identify, assess and mitigate risk across the AI lifecycle – including data poisoning and theft, and the use of third-party services like ChatGPT. Crucially, they’re based on a “Plan-Do-Check-Act” (PDCA) model, which forces organisations to adopt a mindset of continuous improvement. 

As AI is embedded ever deeper into the fabric of the enterprise, cybersecurity teams must get proactive about tackling threats like data poisoning. In many cases, the alternative is exposing the enterprise to untenable reputational and financial risk. But in some scenarios, like autonomous vehicles and industrial control systems, it could be a matter of life and death.

Sam Peters is Chief Product Officer of IO   –   Image: Ideogram

Source: Cyber Security Intelligence