The OpenAI Preparedness Framework is a safety system OpenAI uses to evaluate advanced AI models for capabilities that could cause severe or catastrophic harm. It measures dangerous capabilities, applies safeguards, and uses defined risk thresholds to help determine whether a model can be deployed or needs stronger protections.
For readers searching for the OpenAI 4 Risk Levels, the original framework used four ratings: Low, Medium, High, and Critical. However, OpenAI has since updated the framework and its operational approach, so it is important to understand both the original four-level scale and how the newer framework works.
This matters when evaluating concerns such as the cybersecurity risks discussed in our broader guide to OpenAI Astra.
What Is the OpenAI Preparedness Framework?
The OpenAI Preparedness Framework is essentially a testing and governance system for frontier AI models.
Instead of asking only whether a model produces unsafe answers, OpenAI evaluates whether increasingly capable models could develop abilities that create serious real-world risks. The framework has historically focused on areas such as cybersecurity, chemical/biological/radiological/nuclear (CBRN) threats, persuasion, and model autonomy.
The important distinction is capability.
A model might not be intentionally harmful, yet a powerful enough model could potentially help someone discover software vulnerabilities, provide dangerous biological information, influence people at scale, or perform increasingly autonomous tasks.
That is what the Preparedness Framework is designed to track.
OpenAI 4 Risk Levels Explained
The original scorecard system used four risk levels: Low, Medium, High, and Critical. OpenAI described Low as indicating that the relevant category was not yet a significant problem, while Critical represented the maximum level of concern.
1. Low Risk
Low means testing has not demonstrated capabilities that reach a significant threshold for that particular risk category.
For example, a model could be evaluated for cybersecurity capabilities and still fall below the threshold at which OpenAI considers it capable of meaningfully increasing real-world offensive cyber risk.
Low does not mean “zero risk.”
It means the tested capability remains below the relevant threshold.
2. Medium Risk
Medium represents a more noticeable level of capability.
This is where the model has demonstrated stronger performance in a risk category. However, the framework’s deployment rules have historically still allowed deployment when the post-mitigation rating remained at Medium or below.
GPT-4o provides a useful real example. Its published scorecard rated cybersecurity as Low, CBRN as Low, persuasion as Medium, and model autonomy as Low. Because the highest category determined the overall rating, its overall Preparedness risk was Medium.
So Medium did not automatically mean “do not release.”
It meant the model had crossed a level that required the safety process to pay closer attention to that capability.
3. High Risk
High is where the situation becomes much more serious.
Under the earlier framework, OpenAI stated that models reaching High or Critical risk could not simply be deployed as-is. Mitigations needed to reduce the post-mitigation risk to Medium or below before deployment.
This creates an important distinction between pre-mitigation and post-mitigation scores.
Imagine a model initially tests at High for cybersecurity. OpenAI could then add restrictions, monitoring, training changes, or other safeguards and evaluate it again.
If those measures reduce the score to Medium, the model may satisfy the earlier deployment requirement.
4. Critical Risk
Critical was the highest level in the original four-level scale.
It represented a capability associated with the most severe level of concern. The original framework treated Critical capabilities more restrictively than High ones, including restrictions around continuing development.
This is the level readers should not interpret as simply “a slightly worse High.”
The difference is about the potential severity and nature of the capability.

How the OpenAI Preparedness Framework Scorecard Works
A Preparedness Framework Scorecard gives readers a much easier way to understand the results of these evaluations.
Instead of assigning one vague safety number to an entire AI system, OpenAI has published results by risk category. A model might score differently for cybersecurity, CBRN, persuasion, and autonomy.
The highest risk category determined the overall risk under the earlier framework. GPT-4o demonstrates this clearly: although three categories were Low, persuasion was Medium, resulting in an overall rating of Medium.
That approach is useful because an average score should not obscure one particularly dangerous capability.
Pre-Mitigation vs Post-Mitigation
This is one of the most important details to understand when reading a scorecard.
Pre-mitigation risk asks what the model can do before the relevant safety measures are applied.
Post-mitigation risk asks what the risk looks like after safeguards have been implemented.
A model can therefore have a higher initial capability assessment but a lower final risk rating after mitigation.
OpenAI has also warned that its evaluations should be treated as a lower bound, as additional prompting, fine-tuning, longer interactions, or different scaffolding could elicit capabilities not observed during testing.
What Changed in the New OpenAI Safety Framework?
Here is where many older explanations of the Preparedness Framework become confusing.
According to official OpenAI Preparedness Documentation, the updated safety framework focuses on tracking high-risk capabilities across cybersecurity, CBRN, and autonomous model development. The newer Version 2 framework goes even further: OpenAI explicitly says it is removing the terms Low and Medium from the framework because those levels were not operationally involved in its Preparedness work. It focuses its tracked categories on biological and chemical capability, cybersecurity, and AI self-improvement.
That means an article explaining the “four OpenAI risk levels” should not present Low, Medium, High, and Critical as if they are still the complete current operational system.
They are the historical four-level scale.
The current framework is more threshold-based.
How OpenAI Uses Risk Testing in Practice
The process can be understood in a simple sequence.
First, OpenAI identifies a capability that could create severe harm and determines whether it belongs in a tracked risk category.
Next, researchers evaluate the model using techniques intended to expose its strongest relevant capabilities. These can include specialized testing, prompting, scaffolding, and other capability-elicitation methods.
The results are then reviewed through OpenAI’s safety governance process. If a concerning capability crosses an applicable threshold, OpenAI applies safeguards and evaluates whether those protections sufficiently reduce the associated risk.
The practical question is not simply:
“Is this AI dangerous?”
It is closer to:
“What dangerous capability can this system demonstrate, how severe could the resulting harm be, and are the safeguards strong enough for the level of capability?”
That distinction is especially important for cybersecurity.
A model that can explain basic security concepts is very different from one that can reliably discover and exploit serious vulnerabilities. The framework is intended to detect that jump in capability before deployment decisions are made.
Why the Framework Matters for AI Cybersecurity
Cybersecurity is one of the clearest examples of why preparedness testing matters.
A stronger reasoning model may help legitimate security researchers analyze code, identify vulnerabilities, or automate defensive work. But the same underlying capability could potentially lower the technical barrier for malicious actors.
OpenAI’s published system cards show how these evaluations have been applied to real models. For example, GPT-4o was tested on cybersecurity tasks, including Capture the Flag challenges, while later models have been evaluated using newer versions of the Preparedness Framework.
This is why a model’s general intelligence score tells you very little about its frontier safety risk.
You need to know what it can actually do in high-risk domains.
What Should You Look For in an OpenAI Risk Scorecard?
When you read a Preparedness result, don’t look only at the headline rating.
Check which risk category received the highest designation. Then determine whether the number refers to pre-mitigation or post-mitigation capability. Finally, check which version of the Preparedness Framework was used.
That last step is easy to miss.
A “Medium” result from an older scorecard and a “High capability” designation under the newer framework are not necessarily describing the same measurement system.
For anyone researching AI safety, this is one of the easiest ways to avoid misreading OpenAI’s published evaluations.
Frequently Asked Questions
What are the four OpenAI risk levels?
The original OpenAI Preparedness Framework used four levels: Low, Medium, High, and Critical. Low represented the lower end of the risk scale, while Critical represented the highest level of concern. OpenAI’s newer framework has moved away from using Low and Medium as operational Preparedness levels.
Can an AI model with High risk still be deployed?
Under the earlier framework, a model reaching High risk could not be deployed until mitigations reduced its post-mitigation risk to Medium or below. The updated framework instead focuses on whether safeguards sufficiently minimize the severe risks associated with High or Critical capabilities.
Does a Low risk score mean an AI model is completely safe?
No. A Low rating applies to a particular evaluated risk category and threshold. It does not mean the model has zero risk, nor does it guarantee that every possible harmful capability has been identified. OpenAI itself describes Preparedness evaluations as a lower bound on potential risk.
The key takeaway is simple: the OpenAI Preparedness Framework is not just a four-color safety label. It is an evolving evaluation system designed to identify dangerous frontier capabilities, measure how serious they could become, and connect those findings to safeguards and deployment decisions.
And because OpenAI continues to update the framework as models become more capable, always check which version of the framework a scorecard uses before comparing risk ratings across different AI systems. For a practical example of why these assessments matter, continue with our guide to OpenAI Astra and the cybersecurity risks associated with increasingly capable AI systems.







