Skip to main content

From bugs to bypasses: adapting vulnerability disclosure for AI safeguards

Exploring how far cyber security approaches can help mitigate risks in generative AI systems

,
Visual Generation via Getty Images

As AI systems become more powerful, so do the risks of misuse. From the tangible real-world harms caused by insufficient safeguards which we see today, to the longer term potential risks from malicious actors, the stakes are high. This blog explores how traditional cyber security practices can help mitigate these risks, with a particular focus on public disclosure programmes.

This blog is aimed at:

  • decision-makers involved in the design, development, and deployment of AI systems
  • researchers exploring AI safety and security
  • anyone interested in frontier AI and/or cyber security





There are additional potential benefits to public disclosure programmes, such as:

  • Encouraging a culture of responsible disclosure, by incentivising and promoting ethical behaviour (although there are some caveats, see below), and encouraging collaboration between researchers and developers.

  • Increasing brand awareness and engagement among the security community and projecting a sense of security around the product.

  • Providing an opportunity for researchers to practice and demonstrate a range of real-world security skills.

Other factors to consider:

  • Companies don’t necessarily need to provide a financial incentive to get many of the benefits. It's likely beneficial to provide a range of incentives on top of purely financial ones.

  • The breadth and diversity of evaluation of public programmes should supplement, not replace, deeper security evaluations.

  • There are significant overheads associated with triaging and managing reports.

  • It won’t be effective unless the developers have good foundational security practices in place.


Please note: We welcome programmes like SBDPs and SBBPs to encourage and support cyber security analysis of AI models. But note that the presence of an SBDP and SBBP does not automatically mean the model or system is safe or secure. We’re encouraging further research on this and other questions (see below).



Written by

Dr Kate S Technical Director for Security of AI Research, NCSC
Dr Robert Kirk Research Scientist, Safeguards, DSIT