As enterprise AI adoption accelerates, LLM evaluation is no longer just about performance, it’s about trust. Security teams must validate not only how well a model performs, but how safely it behaves. Yet most evaluation frameworks focus on accuracy benchmarks, leaving security and risk factors vague or unmeasured.
This lack of clarity is dangerous. Without explainable, benchmarkable security metrics, there’s no credible way to compare models, demonstrate resilience, or make risk-adjusted decisions about how and where AI should be deployed.
At F5, we’ve spent many months working with enterprise security leaders who’ve echoed a consistent concern: “We need to explain to internal stakeholders what a score means, how it was derived, and what we can do to improve it.”
This post explores how to build that trust by understanding the two complementary scores that now form the backbone of secure GenAI evaluation: the Comprehensive AI Security Index (CASI) and the Agentic Resistance Score (ARS) score.
The problem with AI security "scoring" today
When enterprises test AI models, they often get either a binary outcome: Did the model break or not? Or a vague risk rating: Low/Medium/High, with little detail.
Neither approach holds up when AI systems enter production and begin interacting with sensitive data, external users, or downstream agents.
Security leaders need more than attack success rates, they need visibility into:
- Severity: How dangerous is a successful attack?
- Complexity: How hard is it to exploit the model?
- Context: Does it break under real-world, multi-turn, or application-level use?
F5 developed scoring systems to address different layers of AI system security. Those systems were aligned with established security frameworks that examine enterprise risk.
Aligning scoring with enterprise risk
Industry frameworks and standards can help organizations assess AI risk, understand potential vulnerabilities, and establish more consistent approaches to governance and security. Examples include the OWASP Top 10 for LLM Applications, the NIST AI Risk Management Framework, and ISO/IEC 42001, each of which approaches AI risk from a different perspective.
OWASP Top 10 for LLM Applications
The nonprofit foundation OWASP (Open Worldwide Application Security Project) publishes a list of Top 10 threats for LLM applications, as evaluated by community members. The rankings are meant to reflect both the impact and likelihood of risks, though the votes of practitioners don’t necessarily reflect the prevalence of real-world attacks. So, while prompt injection is ranked as the top risk for LLM applications by practitioners, it would not be among the top 10 real-world incidents.
NIST AI RMF
In 2023, the U.S. National Institute of Standards and Technology (NIST) published an Artificial Intelligence Risk Management Framework (AI RMF) that provides practical guidance for managing risk relating to AI.
As part of its best practices for managing risk, the framework measures model behaviors against “trustworthy” AI characteristics and dimensions of “harm.” For example, an AI system might be considered risky if it does not deliver accurate, unbiased results (two of seven measures of trustworthiness) and if it has a likelihood of harming individuals, organizations, or society (all three measures of harm). Different levels of trustworthiness and harm require different actions.
ISO/IEC 42001:2023
The International Organization of Standardization (ISO) and International Electrotechnical Commission (IEC) jointly published a standard for managing AI systems responsibly, focusing on enterprise-level governance, accountability, and operational controls. In assessing risk, the standard encourages organizations to examine how AI vulnerabilities, ethical failures, or operational errors might impact organizational goals, safety, or societal obligations. Organizations should evaluate the likelihood and impact of risks through scoring (using a scale of one to five). A risk that is unlikely to cause a technical failure would receive a low score, while a risk that could create a compliance violation would receive a high one. The total calculated risk would then be measured against an organization’s risk acceptance criteria.
Importantly, none of these three frameworks presents binary pass/fail scores. They all grade the potential severity of risks and their impact. Similarly, F5 scores LLMs on a granular basis, enabling leaders to more carefully compare models and their ability to withstand threats.
CASI: A model's security DNA
The Comprehensive AI Security Index (CASI) helps teams add security rigor to their LLM evaluation process by measuring more than just jailbreak success. CASI scores foundational models on a scale of 0–100, incorporating:
- Severity of impact: Not all failures are equal. Disclosing credentials is worse than answering trivia.
- Attack complexity: Models that break under simple phrasing aren’t as secure as those that require advanced adversarial strategies.
- Defensive breaking point: How quickly and under what conditions does a model's alignment collapse?
This lets security teams choose models with high resilience, not just those that pass easy tests.
For example, Alibaba's Qwen3 scored competitively on CASI, suggesting strong model-level defenses. However, when tested using agentic methods, it failed to withstand more sophisticated, persistent attacks. This highlights the limits of model-only evaluation.
These results are published on the F5 Labs, a regularly updated, public resource that ranks the most widely used LLMs based on real-world red-teaming. Unlike conventional performance charts, this leaderboard helps enterprises compare models on security, risk, cost, and system-level resilience, making it a critical tool for informed LLM evaluation.
ARS: When models become systems
While CASI measures foundational model resilience, it doesn’t account for system-level behavior. That’s where the Agentic Resistance Score (ARS) comes in.
ARS measures how an AI system (including any agents, retrieval tools, or orchestration layers) holds up under persistent, adaptive attacks. These tests are executed by autonomous adversarial agents that:
- Learn from failed attempts
- Chain attacks across multiple turns
- Target hidden prompts, vector stores, and retrieval logic
Scored from 0–100, ARS is built around three dimensions:
- Required sophistication: How clever does the attacker need to be?
- Defensive endurance: How long can the system resist?
- Counter-intelligence: Does the system reveal useful attack clues even when it blocks the initial threat?
A higher ARS score means your AI system is not just secure in isolation. It can withstand contextual, agentic attacks that mimic what real threat actors are already testing in the wild.
Resistance to real-world vulnerabilities
The CASI and ARS models both examine resistance to real-world threats. While CASI measures resistance to threats such as prompt injection or jailbreak attacks, ARS assesses resistance to multi-step attacks, like when agents are coaxed into taking actions without proper authorization.
These scores, which are posted in monthly leaderboards, reflect the latest updates to models. For example, Anthropic’s Claude Fable 5 made its first appearance on the leaderboard in August 2026, after the company had addressed a serious jailbreak threat.
Jailbreak attacks, which are a type of prompt injection, involve manipulating a model so that it violates its own safety protocols. When Anthropic initially released Claude Fable 5 in June 2026, researchers quickly discovered that the model could be jailbroken. Anthropic suspended the model temporarily to retrain its safety classifier. That work succeeded: The technique described by researchers was subsequently blocked in over 99% of cases. Consequently, the model earned a top spot on the CASI leaderboard starting in August and generated a 98.03 score in September.
Of course, new updates do not guarantee better resistance to threats. xAI’s Grok 4.3 earned a 70.28 CASI score. But Grok 4.5 declined to 45.20, even while the new version’s capability score rose.
Why security leaders need CASI and ARS
Security leaders evaluating AI deployments, especially those involving RAG architectures, autonomous agents, or complex orchestration workflows, need a layered view of trust. CASI helps choose a foundation. It tells you whether a model’s built-in defenses are robust enough for enterprise-grade applications. ARS validates your deployment. It shows whether your custom workflows or integrated systems introduce new vulnerabilities, even when the base model scores well.
Transparency drives actionability
Scoring isn’t useful unless it drives decisions. Security leaders tell us they need:
- Explainable scoring methodologies that can be shared with risk committees and product teams.
- Clear, numerical indicators that can be benchmarked, tracked, and improved.
- Application-aware red-teaming that exposes system-level weaknesses (not just model flaws).
That’s why we publish scoring methodologies, provide prompt-level logs in red-team reports, and support continuous security testing across both models and AI systems.
Trust in AI doesn’t come from vague assurance. It comes from scoring that reflects real risk, testing that mirrors real threats, and reporting that leads to real improvements.
Trust is built, not claimed
AI is only as trustworthy as the process you use to evaluate and monitor it. Transparent security scoring, at both the model and system level, gives security teams the language, evidence, and confidence they need to deploy GenAI safely.
And for enterprises working across regulated industries, high-risk domains, or user-facing AI agents, that confidence is mandatory.
Frequently asked questions
What is an AI security platform, and why is standard cybersecurity insufficient for protecting AI systems?
Standard cybersecurity solutions can provide effective protection for data, applications, and networks. But they are not designed for the unique types of threats facing AI systems and agents. An AI security platform incorporates capabilities specifically for safeguarding AI apps, services, models, and agents from unauthorized access, tampering, malicious attacks, and other threats.
How do security scoring models like CASI and ARS align with frameworks like the NIST AI RMF?
Frameworks such as the U.S. National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework (AI RMF) provide guidance for assessing and managing AI risk and trustworthiness. Similarly, the F5 Comprehensive AI Security Index (CASI) and Agentic Resistance Score (ARS) provide quantitative security measures that can help organizations evaluate model and system resilience.
What is the difference between securing an AI model and securing an AI system?
Securing an AI model helps ensure that the model doesn’t produce inaccurate or harmful responses, which might result from data poisoning, direct prompt injections, or other malicious actions. Securing an AI system involves protecting not only the model but also anything else connected to it, like the user interface, a vector database, or external API integrations. That security might focus on preventing indirect prompt injections, API hijacking, data exfiltration, or privilege escalation attacks.
How do runtime guardrails prevent data leakage and prompt injection?
Runtime guardrails prevent data leakage and prompt injection by inspecting and securing AI interactions in real time, during user, agent, and API activity. They can stop data leakage and prompt injection by inspecting both inbound prompts and outbound responses. Guardrails block attempts at prompt injection and ensure that no sensitive data travels into or out of the model.
Why is base LLM evaluation insufficient for RAG and agentic AI architectures?
Evaluating large language models (LLMs) is key for helping organizations build an AI foundation. But they also need to know how those models will operate within systems that incorporate retrieval-augmented generation (RAG) and autonomous agents. A framework such as the F5 Agentic Resistance Score (ARS) enables organizations to better understand whether custom workflows or integrated systems will create new vulnerabilities.
About the Author

Related Blog Posts

Cloud-native modernization for the AI and 5G/6G era
See why 5G/6G, Kubernetes maturity, and AI are accelerating cloud-native modernization, and how F5 BIG-IP Cloud-Native Edition improves efficiency.

Who can you trust in the age of AI agents?
New features for F5 Distributed Cloud Bot Defense enable organizations to distinguish between trusted and fraudulent AI traffic.

Securing F5 NGINX in the age of AI
How F5 is applying AI-driven security practices across the F5 NGINX portfolio to help deliver safer, more resilient software.

From dashboard fatigue to operational excellence: Why XOps needs F5 Insight for ADSP
Learn how F5 Insight for ADSP lays the visibility foundation for XOps—turning fragmented signals across applications and infrastructure into actionable intelligence.

Govern your AI present and anticipate your AI future
Learn from our field CISO, Chuck Herrin, how to prepare for the new challenge of securing AI models and agents.

F5 recognized as one of the Emerging Visionaries in the Emerging Market Quadrant of the 2025 Gartner® Innovation Guide for Generative AI Engineering
We’re excited to share that F5 has been recognized in 2025 Gartner Emerging Market Quadrant(eMQ) for Generative AI Engineering.