In July 2026, the Future of Life Institute published the summer edition of its AI Safety Index, and the conclusion was as clear as it was uncomfortable: no frontier AI lab scored higher than C+. No distinctions. Not even honorable mentions. The top of the class barely passed, and that top spot went to Anthropic, the company that built its brand on the promise of being the safest.
While language models gain new capabilities week after week, the safety infrastructure that should accompany that progress is falling behind. And that is not an opinion: it is the result of 37 indicators evaluated by seven independent experts from Berkeley, Montreal, Wisconsin, and Oxford.
If your organization already uses tools built on these models, the index results are not an academic curiosity. They are the actual state of the ecosystem on which you built your AI stack.
The Exam Nobody Wanted to Take
The AI Safety Index is published twice a year and evaluates leading AI companies across six domains: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information sharing. Each domain has multiple indicators. The maximum score is 4.0, equivalent to an A in the US academic grading system.
The panel that assigns the grades includes figures such as Stuart Russell, a co-founder of the AI field and professor at UC Berkeley, and David Krueger, one of the founders of the UK AI Security Institute. These are not activists with an agenda: they are technical authorities in the field.
Evidence was collected through June 3, 2026. Four companies (Alibaba, xAI, DeepSeek, and Mistral) did not even respond to the index survey. Their grades were built solely from publicly available information.
The Scoreboard
| Company | Grade | Score |
|---|---|---|
| Anthropic | C+ | 2.66 |
| OpenAI | C | 2.28 |
| Google DeepMind | C | 2.01 |
| Meta | D+ | 1.32 |
| Z.ai | D- | 0.88 |
| Alibaba Cloud | D- | 0.87 |
| xAI | F | 0.65 |
| DeepSeek | F | 0.47 |
| Mistral | F | 0.33 |
Three companies received failing grades: one American (xAI), one Chinese (DeepSeek), and one European (Mistral). The index highlights a particularly striking "European dissonance": the European Union leads global AI regulation, yet the top European AI lab finishes last in safety.
What Worries the Panel Most
The grades are already revealing, but the qualitative findings are even more unsettling.
Pause commitments have vanished. Anthropic, OpenAI, Google DeepMind, and Meta had signed commitments to halt development if certain risk thresholds were reached. All of them have since weakened or withdrawn those pledges. The panel describes this as "moving the goalposts" and argues it has "undermined safety frameworks across the board." Stuart Russell was direct: companies plan to release new systems even when it is demonstrably unsafe to do so.
Military AI has become a cross-cutting problem. Between 2024 and 2026, companies that explicitly banned military applications, including Anthropic, OpenAI, Google DeepMind, and Meta, reversed that position and are actively pursuing defense contracts. The panel identified this as an emerging current-harm risk.
Existential safety is the weakest domain industry-wide. No company exceeded C- in this domain. Most score D or below. Dominant paradigms bet on detecting problems. But detection is not prevention.
Safety rhetoric outpaces revealed behavior. At Google DeepMind, OpenAI, and xAI, reassuring public messaging from leadership diverges from commercial conduct and legislative positions those same companies take.
What It Means to Use AI from a Company That Failed
The index is clear that it evaluates companies as organizations, not the specific models or products they deploy. An F from xAI is not a grade for the Grok model; it is an assessment of how xAI manages risk as a company.
But that distinction does not remove the practical implication. When a company fails, the report describes what was concretely missing:
xAI (F): No evidence of a safety team with real weight. Dangerous-capability evaluations have enormous gaps, and no procedure connects those results to deployment decisions. If they detect a serious problem in a model before launch, there is no formal mechanism that forces a pause.
DeepSeek (F): No published safety framework. No visible governance structure. The only real protection is Chinese government regulation, which the panel describes as "complete passivity" in the face of existential risks from advanced AI.
Mistral (F): Leadership actively downplays frontier risks rather than articulating a control strategy. No published safety framework and unacceptable performance on safety benchmarks. And they are the ones building the models deployed in several European corporate products.
If your organization uses any of these models, the question is not whether the AI works. The question is who responds when something goes wrong and what protocol exists to prevent that from reaching your critical workflows.
Can You Trust It, Then?
It depends on what for.
The honest answer is not yes or no: "trusting AI" is the wrong question. The right question is what level of trust is reasonable for what type of task.
For low-risk tasks (drafting a document, summarizing content, generating ideas), models from any lab on the list work fine and the provider's safety level is almost irrelevant. If the model hallucinates, you catch it.
For processes that touch sensitive data, decisions with legal impact, or workflows where an error propagates without human review, the safety level of the lab behind the model starts to matter a great deal. And there is an additional trap: part of the index's indicators are based on standardized benchmarks, which, as we have covered in our benchmark article, can be inflated or saturated in ways that mask real capability gaps. The index compensates with external evaluations and governance indicators, but the limitation is worth keeping in mind.
Using Grok, DeepSeek, or Mistral in a critical process is not just betting on a model with weaker safety benchmark performance. It is betting on a company that has not documented what it does when something goes wrong.
That does not automatically disqualify any of them for all uses. But it does shift the burden of proof: if you choose a provider with an F, the responsibility for controls falls entirely on your organization.
The Point
The highest grade was C+. That does not mean AI is unusable. It means the default optimism when adopting it for critical processes deserves to be replaced by concrete evaluation criteria. Three questions that should be part of any AI vendor assessment:
What happens when the model hits a risk threshold? Is there a formal procedure, or does it depend on an executive's judgment?
Who can stop a deployment? Is there a body with real authority, or just a recommendation?
Do the safety commitments have attached conditions? If the promise to pause depends on competitors doing the same, it is not a commitment.
The index provides the criteria. Applying them is each organization's work. If you want to assess how this applies to your current AI stack, reach out to macareno.net and we will work through it together.
Sources
- AI Safety Index — Summer 2026 — Future of Life Institute
- The Latest AI Safety Rankings Are In. Nobody Gets an A — TIME
