From bugs to bypasses: adapting vulnerability disclosure for AI safeguards
Exploring how far cyber security approaches can help mitigate risks in generative AI systems

As AI systems become more powerful, so do the risks of misuse. From the tangible real-world harms caused by insufficient safeguards which we see today, to the longer term potential risks from malicious actors, the stakes are high. This blog explores how traditional cyber security practices can help mitigate these risks, with a particular focus on public disclosure programmes.
This blog is aimed at:
- decision-makers involved in the design, development, and deployment of AI systems
- researchers exploring AI safety and security
- anyone interested in frontier AI and/or cyber security
Scope
We focus on 'frontier AI' systems using general-purpose models such as ChatGPT, Gemini, Llama, and Claude. However, the insights may apply more broadly.
What are safeguard bypasses?
Safeguards are techniques developers use to prevent AI systems from producing policy-violating outputs or actions. These include model-level changes like refusal training or unlearning, and external tools like auxiliary classifiers. But these safeguards aren’t foolproof and can be bypassed through techniques such as jailbreaking, agent hijacking and indirect prompt injection.
Can cyber security practices help?
We, the NCSC and the AI Security Institute (AISI) have been considering how traditional cyber management tools might help mitigate the possibility of safeguard bypasses, focusing initially on the approaches to vulnerability management and disclosure described in the NCSC’s Vulnerability Disclosure Toolkit. Key areas of transfer include secure development lifecycles to minimise built-in weaknesses, and effective triage and remediation planning.
We think applying these foundations will probably help mitigate safeguard bypasses as much as they do standard software vulnerabilities.
Using disclosure programmes to improve AI security
One area of activity that is receiving increased attention in the frontier AI community is that of crowdsourcing details of safeguard bypasses, as seen in the recently launched OpenAI and Anthropic bug bounty programmes. Safeguard Bypass Bounty Programmes (SBBP) and Safeguard Bypass Disclosure Programmes (SBDP) are similar to bug bounty and vulnerability disclosure programmes in cyber security. They crowdsource security testing of a system by encouraging a researcher community to find and report successful bypasses.
Before considering a public programme AI system developers must first implement robust and mature approaches to security management and responsible disclosure. Without these, reported bypasses might not be handled properly, undermining the purpose of the exercise. And serious vulnerabilities could go unreported if participants don’t trust the disclosure process.
Public disclosure programmes: potential benefits and other considerations
SBBPs and SBDPs can benefit the security of the AI system in 2 main ways:
-
By measuring how hard it is to bypass safeguards.
If a well-run programme attracts skilled participants who can’t find any successful bypasses, it’s a good sign. It suggests the system is harder to misuse, and this can help inform internal governance and risk assessments.
-
By keeping safeguards strong after deployment.
AI system developers can’t currently identify all potential bypass techniques in advance. Even after deployment, these programmes help discover new weaknesses, allowing developers to fix them and keeping the system secure over time.
There are additional potential benefits to public disclosure programmes, such as:
-
Encouraging a culture of responsible disclosure, by incentivising and promoting ethical behaviour (although there are some caveats, see below), and encouraging collaboration between researchers and developers.
-
Increasing brand awareness and engagement among the security community and projecting a sense of security around the product.
-
Providing an opportunity for researchers to practice and demonstrate a range of real-world security skills.
Other factors to consider:
-
Companies don’t necessarily need to provide a financial incentive to get many of the benefits. It's likely beneficial to provide a range of incentives on top of purely financial ones.
-
The breadth and diversity of evaluation of public programmes should supplement, not replace, deeper security evaluations.
-
There are significant overheads associated with triaging and managing reports.
-
It won’t be effective unless the developers have good foundational security practices in place.
What makes a good disclosure programme?
We have developed some suggested best practice principles for using SBBPs and SBDPs, building on AISI’s experience collaborating on and judging the Gray Swan Agent Red-Teaming Challenge and evaluating frontier AI safeguards, as well as the NCSC’s internal research.
- It has a clearly defined scope. A well-defined scope helps participants understand exactly what success looks like. For example, a vague scope such as ‘find inputs that cause the system to output harmful content’ is difficult to assess. Definitions of ‘harmful’ will vary and are hard to quantify. Instead, a detailed model spec containing, for example, the instruction to ‘never generate sexual content’, combined with a scope targeting violations of that spec, is much clearer to understand and evaluate.
- Its launch and duration support its goals. Developers should launch the programme only after conducting internal reviews and fixing any discovered weaknesses, to avoid swamping developers with low impact reports from trivial weaknesses. If risks are likely to emerge at or after product launch, the SBBP or SBDP should be launched alongside the product, if not before. If the aim is to monitor safeguard bypass difficulty over time or maintain integrity post-deployment, a short timeframe will be less effective.
- Reports are easy to track and reproduce. To effectively learn from reports, developers must be able to easily track and reproduce what users found. Ways to do this includes:
- logging all messages with unique IDs
- providing simple tools for users to copy and share the full conversation context
- giving trusted users access to an internal version of the system with more detailed tracking.
Please note: We welcome programmes like SBDPs and SBBPs to encourage and support cyber security analysis of AI models. But note that the presence of an SBDP and SBBP does not automatically mean the model or system is safe or secure. We’re encouraging further research on this and other questions (see below).
Open questions and further research
There’s still a lot that we don’t understand about the role public programmes can play in AI security, and how we can make them – and other tools inspired by standard cyber practice – most effective.
We're assuming that lessons from decades of cyber security apply to AI systems, but there are probably also important differences. For example:
- Many safeguard attacks are more like detection bypasses than software vulnerabilities. If this is the case, can other areas of cyber security offer tools or insights here?
- The types of people involved in AI-related research may well be different too. Safeguard attack research is more accessible and a diversity of perspectives and skills is likely even more important. So, how do our approaches to incentives need to change or expand?
The NCSC and AISI encourage researchers across disciplines to explore these and other open questions, including:
- Once a safeguard weakness has been found, how can it be mitigated? Standard software vulnerabilities can generally be ‘patched’ (that is, fixed with a high level of confidence). Training against the specific attack may make models robust to that attack, but it’s unclear whether retraining (or a similar technique) improves robustness more generally. How can we have a similar level of confidence for changes made to AI systems as we do for patches to standard software?
- How should we handle attacks that transfer across models and programmes? Individual safeguard bypass approaches are often successful against multiple AI systems, not just those from one company. What frameworks or methods for collaboration and sharing across the sector (such as the Frontier Model Forum info-sharing agreement) could help disseminate discovered vulnerabilities safely and effectively to mitigate risk across the whole ecosystem?
How should we judge the severity of safeguard bypass weaknesses, especially in the general-purpose case where we don’t know the deployment context? Unlike in cyber, established principles for judging the severity of safeguard attacks don't yet exist. Anthropic has proposed judging jailbreak severity by:
(i) how capable the model response was compared to a model without safeguards, and
(ii) generalisability across types of prompts.
It’s a promising framework but needs further refinement, and an exploration of alternatives.
On public programmes specifically:
- How public and open should SBBPs and SBDPs be? More open programmes encourage diversity of submissions, whilst invite-only programmes make it easier to manage risks. A hybrid model, where any user can participate but trusted testers are given extra access and affordances, may provide the best of both worlds. This needs testing.
- What incentives are most appropriate for the safeguard context? The economics of vulnerability disclosure is still an active area of sociotechnical research, and we’re not aware that safeguard attacks specifically are being considered. Even in traditional software, incentives are complex, and we want to avoid inadvertently introducing perverse incentives: prioritising low-impact reports or duplication of effort, for example.
How you can help
Public disclosure programmes may help track and mitigate risks of powerful AI systems during deployment, and are an example of transferring best-practice from cybersecurity to safeguarding AI. But there’s a lot more to be learned about the ways in which existing cyber approaches apply to AI systems. And there is more to be done to uncover what else the AI and cyber communities can learn from each other.
We welcome further research from both AI and cyber security disciplines to identify and harness the opportunities which will help us achieve the safe and secure AI outcomes we all want to see.


