No AI company on Earth currently earns better than a C+ on independent safety grading. That is the headline finding of the Future of Life Institute's AI Safety Index Summer 2026, published July 7, 2026, which graded nine of the world's leading AI labs across six domains of safety and security practice. The result is not a single lab failing an easy test. It is an entire industry clustering in the C-to-F range on a scale where the report's own reviewers say the top grade should be far higher by now.
This piece walks through what the index actually measures, how the nine labs scored, where the industry's real weaknesses sit, and what a grade like this can and cannot tell an organization deciding which model to build on.
What the AI Safety Index measures
The AI Safety Index is not a benchmark of model capability. It does not test how well a chatbot answers questions or writes code. It is a governance and policy audit, built by the Future of Life Institute (FLI), the nonprofit behind several prior AI safety open letters, and reviewed by an independent panel of AI safety researchers and policy experts. The Summer 2026 edition evaluated nine companies against roughly three dozen indicators grouped into six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance and Accountability, and Information Sharing.
Anthropic
Meta
xAI
DeepSeek
Mistral AI
Z.aiIn practice, that means the index scores things like whether a lab publishes a formal risk-assessment framework, whether it discloses model evaluations before deployment, whether it has committed to third-party audits, and whether its public safety pledges hold up against what it has actually done. It is closer to a corporate governance rating than a product test. That distinction matters for reading the grades correctly, and we return to it below.

The grades: nine labs, no one above a C+
Anthropic received the highest overall grade in this edition, a C+ with a numeric score of 2.66 on FLI's 4.0-point scale, the highest grade the index has awarded any lab to date. OpenAI and Google DeepMind followed with a C each (2.28 and 2.01 respectively). Meta improved to a D+ (1.32), the only lab to meaningfully rise since the prior edition. Z.ai and Alibaba Cloud both landed at D- (0.88 and 0.87). Three labs received failing grades: xAI (0.65), DeepSeek (0.47), and Mistral (0.33).
Nine labs, nine grades, no one above a C+
Future of Life Institute, AI Safety Index Summer 2026 (published July 7, 2026). Scores on FLI's 4.0-point scale, converted from letter grades. Higher is better.
Anthropic's C+ (2.66/4.0) is the highest grade any lab has ever received in the index. xAI, DeepSeek and Mistral received failing grades (F).
Anthropic led five of the six graded domains, including Current Harms, Safety Frameworks, Governance and Accountability, and Information Sharing, where its published transparency practices and safety research output scored highest among peers. OpenAI led only one domain, Risk Assessment, where its pre-deployment evaluation disclosures scored best. Mistral has publicly disputed part of the methodology, arguing that it structurally penalizes open-weight model releases, since much of what happens to an open-weight model after release is controlled by the deploying organization rather than the lab that trained it. That is a legitimate methodological question, and it is worth keeping in mind when reading Mistral's F grade alongside the closed-weight labs above it.
Where the industry is weakest
Anthropic led five of six graded domains, but no domain scored well
Highest grade any lab achieved in each domain. Future of Life Institute, AI Safety Index Summer 2026.
Green bars are domains Anthropic led (five of six); the gray bar is Risk Assessment, led by OpenAI. Existential Safety was the weakest domain industry-wide, with no lab scoring above D+.
The domain breakdown reveals something more concerning than the overall ranking. Existential Safety, the category covering a lab's preparedness for the most severe failure modes of increasingly capable systems, was the weakest domain across the entire industry. No lab scored better than a D+ in this domain, meaning even Anthropic and OpenAI, the two leaders, were rated as having no credible plan for handling risks at the frontier of capability. Reviewer commentary on the report has been blunt about this gap, with at least one panel member describing the absence of credible existential-safety plans across the industry as a serious failure of the sector as a whole.
The retreat from earlier pledges
Even industry leaders in safety practices are retreating from prior commitments.
Future of Life Institute, AI Safety Index Summer 2026
The report's most striking theme is not the low grades themselves but the direction of travel. Several labs, including Anthropic, OpenAI, Google DeepMind, and Meta, had previously committed to pausing development of a model unilaterally if it approached specified risk thresholds. According to reporting on the index, all four have since weakened or walked back those pledges, in some cases reframing a pause as contingent on what competitors do first rather than an independent commitment. Anthropic itself withdrew an earlier pledge not to train systems without advance assurance that its safety measures were sufficient. Independent reviewers characterized this pattern as moving the goalposts, warning that it undermines the credibility of voluntary safety frameworks across the industry, even as the same companies compete to ship more capable models faster.
That combination, rising capability paired with softening public commitments, is the core tension the index is trying to make visible. A safety framework that a lab can quietly revise whenever a deadline becomes inconvenient is a weaker signal than the same framework enforced by an external party, which is part of why FLI and other observers continue to argue for independent verification rather than self-reported compliance.
What a grade like this does and does not tell you
It is worth being precise about the limits of this exercise, because a C+ is easy to misread in either direction. The index grades company-level policy, governance structure, and public disclosure. It does not test a specific deployed model, a specific product configuration, or a specific customer contract. A C+ overall grade says nothing about whether a particular model, running under a particular set of safeguards for a particular use case, meets an individual organization's own risk tolerance. Two deployments of the same underlying model can carry very different real-world risk depending on what guardrails, monitoring, and human oversight sit around it.
The methodology also has known blind spots the report itself acknowledges: it relies substantially on what labs choose to publish, it cannot fully audit undisclosed internal practices, and as Mistral's objection illustrates, it treats open-weight and closed-weight release models somewhat differently even though the risk profile of each depends heavily on downstream use. None of that invalidates the exercise. A repeated, methodical, externally reviewed audit of safety governance across the industry is valuable precisely because so little else like it exists. But the honest reading of a C+ is "better governance disclosure than its peers," not "safe to deploy without further diligence."
The practical takeaway
For any organization choosing among AI providers, the index is best used as one input alongside deployment-level testing, not a substitute for it. Safety posture varies by lab, by model, and by how a system is actually configured in production, which is exactly why teams increasingly want the flexibility to route work to different providers rather than lock into a single one. Platforms like Metir AI give teams model-agnostic access to leading providers side by side, so choosing a model on its safety posture, cost, or capability for a given task does not require re-platforming every time a new report comes out.
The Summer 2026 index will not be the last word on this question. FLI has published the index on a roughly semiannual cadence, and the real story to watch is whether the next edition shows labs re-committing to the pledges this one found them retreating from, or whether the gap between stated safety frameworks and demonstrated practice keeps widening as capability keeps climbing.
Sources:
- AI Safety Index, Summer 2026 | Future of Life Institute
- AI Safety Grades Are In: No Lab Tops C+, and the Best Ones Are Retreating | Tech Times
- The 2026 AI Safety Index: Nine AI Labs Graded | The Median
- FLI AI Safety Index 2026: A Buyer's Guide to the C+ Grades | Digital Applied
- Global Big Tech Retreats on AI Safety Pledges, Experts Warn | Seoul Economic Daily
Image credits
Header image and in-body portrait: Max Tegmark, president and co-founder of the Future of Life Institute, photographed at Web Summit in Lisbon in 2024. Both photos by Web Summit via Wikimedia Commons (full frame, cropped portrait), licensed under CC BY 2.0. Both photographs were taken in 2024, before the Summer 2026 index was published; they depict Tegmark as an individual and his FLI affiliation, not the specific report launch.
