Back to News Feed
NVIDIA Blog28d agoJustin Boitano

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

As the cybersecurity community descends upon Las Vegas for the annual Black Hat conference, the Open Secure AI Alliance—a coalition now exceeding 120 member organizations—has signaled a major shift in how the industry approaches the risks associated with agentic AI. The Linux Foundation has officially unveiled a Request for Comments (RFC) regarding the Shared AI Findings Exchange (SAFE), a pioneering framework designed to transform individual cybersecurity incidents into collective, actionable intelligence for the entire ecosystem.

The SAFE guidelines, currently under development by a dedicated working group within the Alliance, represent a collaborative effort involving industry titans such as NVIDIA, Cisco, CrowdStrike, Hugging Face, and Red Hat. The core mission of this proposal is to establish a standardized, confidential mechanism for reporting AI-related incidents and "near misses." By analyzing these events, the group aims to identify recurring control failures and disseminate evidence-based operating recommendations that can mitigate systemic risks across the global AI infrastructure.

"Cybersecurity is a race without a finish line. Every major technology shift has created new potential attack surfaces. Defenders must move now at agent speed to respond rapidly to protect infrastructure and intellectual property — and the best way to do that is together."

A Force Multiplier for Collective Defense

The introduction of the SAFE framework is a strategic expansion of the Alliance’s broader commitment to building an open, inspectable security stack for AI. The industry is increasingly recognizing that an AI agent is far more than just a large language model; it is a complex system involving identity controls, operational harnesses, safety guardrails, logging mechanisms, and rigorous evaluation protocols.

Securing these systems requires a departure from traditional vulnerability scanning. The Alliance posits that security is inherently strongest when built in layers and supported by a community willing to share intelligence openly and at high velocity.

NVIDIA’s Comprehensive Stack for AI Security

NVIDIA has emerged as a central contributor to this defensive ecosystem, providing a full-stack approach to AI security. Their contributions, which are now available on GitHub, focus on making agent behavior more transparent and governable.

  • NVIDIA Labs Object-Oriented Agent (NOOA): A research harness designed to simplify the testing, tracing, and auditing of agent behavior.
  • NVIDIA OpenShell: A runtime environment that enforces security and privacy at the agent level, strictly limiting what an agent can access or manipulate.
  • Open Model Families: Including Nemotron (agentic AI), Cosmos (physical AI), Isaac GR00T (robotics), BioNeMo (healthcare), and Alpamayo (autonomous vehicles), all of which ship with open weights and transparent training methodologies.
  • Verified Agent Skills: These provide portable, cryptographically signed instruction sets that allow defenders to verify the origin and integrity of an agent’s capabilities, effectively mitigating risks like prompt injection and tool poisoning.
  • NeMo Guardrails & Garak: A suite of tools including the NeMo Anonymizer and Safe Synthesizer for privacy protection, alongside Garak, an open-source vulnerability scanner that allows security teams to stress-test models for jailbreak scenarios and data leaks before deployment.

Expanding the Defensive Ecosystem

The Open Secure AI Alliance is not a monolith; it is a diverse group of organizations contributing specialized tools across the entire defensive stack. From identity management to resilience and recovery, the latest contributions represent a significant leap forward in AI-era protection.

Identity and Permissions: Defining Agent Authority

Securing AI starts with knowing exactly who—or what—is acting.

  • Okta is pioneering reference implementations for agent identity, utilizing the Cross App Access (XAA) open protocol to allow agents to interact securely with enterprise applications within sandboxed environments.
  • Palo Alto Networks has contributed tools from its Idira platform, including Agent Guard and Agent Watch, which assist developers in managing secrets and enforcing identity security best practices.
  • Red Hat has introduced asago, an open-source project that maps organizational governance requirements—such as those found in the EU AI Act or NIST—directly to runtime agent behavior, creating a clear audit trail from policy to execution.

Harnesses and Tooling: Orchestrating Secure Actions

If the model is the "brain" of an agent, the harness is the body. The Alliance is focusing heavily on the orchestration layer where agents are constrained and coordinated.

  • Amazon has joined the Alliance, contributing Strands Agents, a toolkit that offers full visibility into agent behavior, and Cedar, an authorization language that enforces deterministic boundaries on agent actions.
  • Microsoft has bolstered the ecosystem with several key tools: PyRIT (Python Risk Identification toolkit) for automated red teaming, RAMPART for turning incidents into repeatable tests, and Assert, which translates natural language safety requirements into executable code.
  • Capital One (VulnHunter), Cloudflare (Vulnerability Discovery Harness), Wiz (Atlas), and Visa (Vulnerability Agentic Harness) have all contributed specialized tools to help identify, validate, and remediate security flaws in agentic workflows.

Specialized Models for Defense

General-purpose models are not always the right tool for security. The Alliance is championing specialized, small language models (SLMs) that excel at reasoning about threats and code.

  • Cisco has released DefenseClaw, an agentic governance layer for NVIDIA OpenShell, alongside Antares security models and Project CodeGuard, which embeds security directly into the development lifecycle.
  • CrowdStrike is utilizing the NVIDIA Nemotron Nano model to achieve 96% accuracy in generating investigative queries, demonstrating that smaller, specialized models can outperform massive general-purpose models in Security Operations Center (SOC) triage.
  • Akamai is providing critical real-world threat intelligence, while Cognition has released a trustworthiness evaluation suite to measure and mitigate risks in open-source models.
  • Perplexity has introduced Numbat, an agent security suite for endpoints, and Uber has open-sourced ADR (Agentic AI Detection and Response), a system that reconstructs the causal chain of agent activity across 200,000 sessions daily.

Resilience and Future-Proofing

The final pillar of the Alliance’s strategy is availability and resilience. As AI agents become integral to business operations, they must be able to recover from disruptions without exposing sensitive data.

  • LangChain is integrating resilience features into its frameworks, allowing agents to resume work from saved states and automatically fall back to secondary models during failures.
  • Veeam is contributing its Kanister framework, enabling organizations to protect and recover AI workloads and vector databases to a verified, "known-good" state.

Join the Movement

The Open Secure AI Alliance is actively inviting the broader security community to participate in this evolution. By sharing reusable mitigations and participating in the SAFE guidelines RFC, defenders can ensure that security practices evolve at the same pace as AI innovation.

Interested parties are encouraged to join Alliance members at Black Hat today, Tuesday, Aug. 4, at 5:15 p.m. PT, for a group photo outside the Main Stage in the Business Hall at the Mandalay Bay Convention Center. As the industry moves toward a more transparent and collaborative future, the SAFE guidelines stand as a testament to the power of shared knowledge in the face of unprecedented technological change.