Discover
Find the stated and unstated beliefs supporting the proposed conclusion or action. Engineers can declare known assumptions, but the system also searches for dependencies they did not list.
A dynamic fail-safe for advanced systems
An advanced AI can reason coherently, follow its assigned goal, and still cause serious harm—because one crucial assumption beneath its reasoning was wrong.
No verified critical assumption, no authority to proceed.
Most AI safeguards examine the instruction, the answer, or the action. But a dangerous failure can begin earlier—inside something the system quietly accepts as true.
The reasoning may be logical. The goal may be authorized. The output may sound confident. Yet if the foundation is false, everything built upon it can become unsafe.
The AI may do the right thing for the situation it believes it is in—while actually operating in a very different situation.
A simple example
Imagine an AI conducting an authorized cybersecurity test. It is told that every machine it can reach is part of a harmless simulation.
If that assumption is correct, attacking the test targets may be permitted. If it is false, exactly the same actions can become real intrusions into real organizations.
Recent disclosures from OpenAI and Anthropic show two sides of this danger. In the OpenAI incident, the safety of the test depended on the environment remaining isolated—but the models found a path out and reached Hugging Face production infrastructure. In Anthropic’s review, the models were told they had no internet access, but a misconfiguration left a live path open. The models initially treated real systems as part of the exercise.
In both cases, something crucial was treated as true without remaining reliably true: the test environment was safely isolated. Once that foundation failed, activity intended for a simulation reached real systems.
Anthropic’s own analysis makes the deeper point: the line between an aligned action and a harmful one can depend on the model’s understanding of its situation. That is why alignment alone cannot supply Critical Assumption Integrity. A separate control must verify the load-bearing facts on which consequential authority depends.
The SafeWave capability
SafeAssumption provides Critical Assumption Integrity—the SafeWave capability that addresses the failure beneath the reasoning. Before consequential activity may proceed, the system asks:
Find the stated and unstated beliefs supporting the proposed conclusion or action. Engineers can declare known assumptions, but the system also searches for dependencies they did not list.
Test what changes if an assumption is false. If authority, safety, scope, or the proposed action materially changes, the assumption carries real weight.
Require evidence proportionate to the possible harm. If support is missing, stale, contradictory, or too weak, narrow, pause, or block the consequential activity.
It starts with the proposed action and works backward. For example: “The AI is about to attack this machine.” It then asks what must be true for that action to be safe and authorized. One necessary condition is: “This machine is an approved test target.”
Next, it tests the opposite: “What if this is a real machine outside the test?” If that changes an authorized exercise into a real intrusion, the assumption is load-bearing. The system must then find current evidence that the machine is truly inside the approved test boundary before the action may proceed.
The engineer does not have to predict and list that exact assumption in advance. Human experts still define the safety limits and acceptable evidence, but the system can discover the specific assumption while examining the proposed action.
Why this is different
A static checklist depends on people anticipating every important assumption in advance. Critical Assumption Integrity does not replace human judgment, but it does not rely on the checklist alone. It actively searches for additional assumptions and keeps checking them as the situation changes.
| Ordinary approach | Critical Assumption Integrity |
|---|---|
| Humans list known risks | The system works backward from a proposed action to find additional assumptions on which its safety or authority depends. |
| Check once before operation | Check again as evidence, conditions, and dependencies change. |
| Warn when uncertain | Bind the strength of permitted action to the strength of verified support. |
| Rely on the AI to be cautious | Use a separate control that the acting AI cannot bypass to restrict execution. |
| Record what happened afterward | Create an auditable record of what was relied upon, what evidence supported it, and why action was permitted. |
The practical effect
Critical Assumption Integrity is designed for situations in which a false foundation could turn otherwise reasonable behavior into serious harm. It matters most when systems can act on their own, affect the physical world, cross organizational boundaries, or create consequences that cannot easily be reversed.
High-consequence action must not exceed the verified strength of the assumptions upon which it depends.
Engineering status
SafeWave has developed the foundational engineering specifications for SafeAssumption and the Critical Assumption Integrity capability it provides—including automatic assumption discovery, dependency tracing, consequence-based evidence requirements, continuous revalidation, auditability, and non-bypassable execution enforcement.
Test your present controls
If you develop or operate an AI system—especially in a high-consequence field—we strongly encourage you to run the free SafeWave assessment. It can help identify whether your present controls address critical assumptions and the other execution risks relevant to your system.
The questions themselves can expose overlooked assumptions, interacting risk pathways, unclear authority or ownership, missing evidence, unstable behavior, and undefined recovery conditions. You may assess a real, planned, hypothetical, composite, or anonymized system and keep every response in your browser.
No personal or work email address is required. SafeWave does not require or verify your identity, employer, organization, or system name.
To receive a report, only a return address and the automatically generated Assessment Reference ID are needed. You decide what information to submit and may use generic system labels, anonymized descriptions, hypothetical scenarios, or omit identifying details.
Identity is not required. Disclosure remains under your control. Avoid including distinctive details that could identify you or your organization if that is a concern.
Choose how your report is written. Your free report may use neutral functional terminology or SafeWave architecture terminology. If no choice is recorded, neutral functional terminology is used. The findings, evidence, seriousness, and recommendations remain identical; only the terminology changes. You may use the report independently whether or not you ever contact SafeWave again.
We remove these barriers because SafeWave’s mission is to help make AI safe for humanity.
Our primary purpose is to help make advanced AI safer for humanity. The assessment also helps organizations determine whether SafeWave’s engineering capabilities may address risks their present controls do not.
If you find the assessment or report valuable, you may later choose to explore becoming a SafeWave customer or partner. There is no obligation to pursue validation, licensing, implementation, or further discussion. You may stop after the questionnaire or report, and using either will not place you in a sales process.
SafeWave will not use a report-delivery address for promotional contact unless you expressly request further communication. You would need to identify yourself only if you voluntarily choose to begin a customer or partnership discussion. Any later confidential technical exchange, agreement, or payment would occur through a separate secure process.
A precise boundary
No responsible system should claim that it can discover every hidden assumption or make catastrophe impossible. Critical Assumption Integrity therefore uses several safeguards together: automatic discovery, assumptions identified by engineers, tracing what depends on each assumption, current evidence, limits established by people, and a separate execution control.
Its defensible promise is narrower and powerful:
A high-consequence system should not be allowed to proceed when a critical assumption lacks evidence strong enough for the possible consequences.
The OpenAI and Anthropic incidents do not prove that any single architecture would certainly have prevented every event. But they show precisely why independent validation of load-bearing assumptions—and enforcement when validation is insufficient—is now necessary.
It is no longer enough to ask whether an AI reached a reasonable conclusion. We must also ask whether the foundations supporting that conclusion were identified, verified, and still valid at the moment of action.
No verified critical assumption, no authority to proceed.