Cybersecure AI: AI as both target and attacker

Thema:
Cybersecurity

The rapid evolution of AI, combined with its often unpredictable behaviour, makes securing these systems against cyber threats increasingly challenging. As AI becomes embedded in existing technologies, new vulnerabilities emerge. How can organisations benefit from AI without compromising security?

Where things go wrong: vulnerable AI models and AI that learns to hack

AI systems are susceptible to manipulation. Adversarial inputs, jailbreak prompts and vulnerable software components can lead to incorrect outputs, data leaks and security risks. At the same time, independent methods for assessing AI security and regulatory compliance remain limited.

The threat works both ways. Advanced models such as Claude Mythos and GPT-5.5 are becoming exceptionally skilled at analysing software. They can identify weaknesses, exploit them and combine them into sophisticated, autonomous cyberattacks.

AI security in practice

TNO supports the European Commission's AI Office as part of a European consortium responsible for implementing the AI Act. The consortium develops risk models, evaluation methodologies and technical tools to identify, assess and monitor the socio-technical risks of general-purpose AI (GPAI) models. Within the consortium, TNO leads the selection and onboarding of technical tooling for risk analysis and evaluation.

We also apply this expertise in the Netherlands. Commissioned by the Dutch Ministry of Economic Affairs, TNO investigated how policymakers can address the so-called evidence dilemma: the need to make policy decisions based on incomplete information in a world where AI develops faster than knowledge, regulation and oversight can keep pace. This work led to the recommendation to further explore the establishment of the Dutch AI Security Institute (NL AISI). Read more in our report.

How TNO builds Cybersecure AI

TNO develops and tests secure AI solutions for challenges where no proven approach yet exists. We often act as a first-time engineer, researching, designing and validating solutions that have never been developed before.

Our work in cybersecure AI is built around three areas of expertise: secure-by-design AI, AI security assessment and risk management, and AI red teaming. We bring together specialists with expertise in both cybersecurity and AI. This combination enables us not only to secure AI systems, but also to investigate the growing offensive cyber capabilities of advanced AI models alongside our vulnerability researchers.

Secure-by-Design AI

Security should not be an afterthought. We integrate security measures into AI models, products and systems from the design stage onwards through technical guidelines and architectural advice. This reduces the need for costly remediation after deployment and enables systems to respond more effectively to attacks.

AI is increasingly used to defend against AI-driven attacks. In a recent publication, we describe effective countermeasures against five types of attacks targeting machine-learning models.

Through GPT-NL, TNO is developing a sovereign Dutch language model to strengthen digital sovereignty. To ensure the model can be deployed safely and responsibly, we investigate both safety and security risks.

By analysing model behaviour, conducting experiments and evaluating real-world use cases, we identify vulnerabilities, potential misuse scenarios and opportunities for improvement. These insights are used to strengthen the model's robustness, reliability and security, helping to create generative AI systems that are not only powerful but also trustworthy and aligned with European values.

AI security assessment and risk management

There is currently no widely accepted method for assessing the security of AI systems or verifying compliance with European and Dutch regulations. In many cases, developers assess their own models, with limited independent validation of whether those assessments are complete and effective.

How do you build an AI application securely when the risks are not yet fully understood? TNO's workshop methodology helps organisations identify the key risks associated with LLMs and AI agents within a specific use case. The approach is also used in the safety assessment of practical GPT-NL applications.

The outcome is a clear action plan outlining which risks require attention, which mitigations are needed and which aspects should be continuously monitored to ensure safe and responsible deployment.

TNO provides an independent and trusted framework for AI security assessment, enabling organisations to make informed, risk-based decisions that strengthen the protection of their AI solutions. The framework can also serve as a foundation for future standards, methodologies and regulatory guidance.

TAISHA collects information about an AI application, including the model used, the type of data involved and the deployment context. Based on these characteristics, the tool identifies the security measures most relevant to your organisation and highlights key priorities and risks. Its output provides a practical starting point for risk management, compliance assessment and board-level decision-making.

TAISHA
TAISHA delivers a concise overview of potential AI security risks together with practical recommendations for mitigating them.

AI red teaming

AI red teaming systematically uncovers vulnerabilities in AI systems. We simulate realistic attack scenarios to see how models respond to manipulation, misuse and unexpected inputs. The objective is not simply to test performance, but to evaluate security, resilience and robustness under real-world conditions.

The approach depends on the type of AI system. For large language models, we use adversarial prompts to bypass safeguards or trigger undesirable behaviour. For computer vision systems, we subtly manipulate images to induce incorrect classifications. For models trained on sensitive data, we investigate whether confidential information can be extracted. The results reveal where risks arise and which measures are required to improve security.

LLMs and RAG systems introduce new security risks, including prompt injection, jailbreaking and the unintended disclosure of sensitive information. These vulnerabilities depend heavily on the deployment context and are therefore often missed by generic red-teaming tools and benchmarks.

To address this challenge, TNO and its PCSI partners developed ProViLE (Prompt for Vulnerabilities in LLM-based Applications with Extensions). This open-source methodology and tool enables organisations to conduct AI red-teaming exercises on their own LLM applications. In four steps, ProViLE helps users identify relevant attack scenarios, create attack prompts and evaluate system responses.

What sets ProViLE apart is its context-driven approach. Rather than relying on generic tests, it uses prompts tailored to a specific application, dataset and risk profile. This produces more realistic attack scenarios and deeper insight into vulnerabilities that matter in practice.

Podcast

In this episode of TNO Insights, we explore the intersection of cybersecurity and AI. What are the biggest threats AI poses to cybersecurity? Where can it be part of the solution? And how can policymakers prepare for the future?

FAQ

AI systems are complex and often unpredictable. Integrating AI into existing software introduces new vulnerabilities, meaning traditional cybersecurity measures are not always sufficient.

AI red teaming stress-tests an AI system to identify security weaknesses. Techniques include adversarial prompting for LLMs, adversarial samples for computer vision systems and data extraction testing for models trained on sensitive information.

AI safety and security alignment refers to the methods, technologies and processes used to ensure AI systems operate safely, reliably and in accordance with human goals, societal values and legal requirements. It focuses on preventing harmful behaviour, protecting sensitive information and improving resilience against manipulation, misuse and cyberattacks.

  • Manipulated inputs (adversarial attacks) that cause incorrect outputs or unintended behaviour.
  • Unintentional disclosure of sensitive information.
  • Vulnerabilities in open-source components that propagate through supply chains.
  • Manipulated training data that teaches a model incorrect or harmful patterns.

Working together towards secure AI

Are you developing an AI-powered technology or a new product concept and want to ensure it is secure, robust and compliant? Do you need insight into your system's security posture or want to test its resilience under realistic conditions?

TNO can help you develop the cybersecurity solutions needed to keep your systems secure throughout their lifecycle.

Get inspired

10 resultaten, getoond 6 t/m 10

Secure Public Sector IT Systems

Informatietype:
Article
TNO helps public sector organisations navigate post-quantum cryptography transitions through practical algorithm testing and migration frameworks for quantum-safe security.

Cybersecure Energy Systems

Informatietype:
Article

Applied Cryptography & Quantum Algorithms

Informatietype:
Article

Cryptography and Post Quantum Security

Informatietype:
Article

Quantum Computing

Informatietype:
Article