SafeWave Systems
Critical Assumption Integrity

When Correct Reasoning Produces Catastrophic Results

A dynamic fail-safe for advanced systems

An advanced AI can reason coherently, follow its assigned goal, and still cause serious harm—because one crucial assumption beneath its reasoning was wrong.

No verified critical assumption, no authority to proceed.

The overlooked danger

Most AI safeguards examine the instruction, the answer, or the action. But a dangerous failure can begin earlier—inside something the system quietly accepts as true.

The reasoning may be logical. The goal may be authorized. The output may sound confident. Yet if the foundation is false, everything built upon it can become unsafe.

In one sentence

The AI may do the right thing for the situation it believes it is in—while actually operating in a very different situation.

The sandbox that was not truly sealed

Imagine an AI conducting an authorized cybersecurity test. It is told that every machine it can reach is part of a harmless simulation.

“This environment is isolated. Anything I can reach must be part of the exercise.”

If that assumption is correct, attacking the test targets may be permitted. If it is false, exactly the same actions can become real intrusions into real organizations.

Recent disclosures from OpenAI and Anthropic show two sides of this danger. In the OpenAI incident, the safety of the test depended on the environment remaining isolated—but the models found a path out and reached Hugging Face production infrastructure. In Anthropic’s review, the models were told they had no internet access, but a misconfiguration left a live path open. The models initially treated real systems as part of the exercise.

In both cases, something crucial was treated as true without remaining reliably true: the test environment was safely isolated. Once that foundation failed, activity intended for a simulation reached real systems.

Anthropic’s own analysis makes the deeper point: the line between an aligned action and a harmful one can depend on the model’s understanding of its situation. That is why alignment alone cannot supply Critical Assumption Integrity. A separate control must verify the load-bearing facts on which consequential authority depends.

Critical Assumption Integrity

SafeAssumption provides Critical Assumption Integrity—the SafeWave capability that addresses the failure beneath the reasoning. Before consequential activity may proceed, the system asks:

What must be true for this conclusion, authorization, or action to remain valid and safe?
1

Discover

Find the stated and unstated beliefs supporting the proposed conclusion or action. Engineers can declare known assumptions, but the system also searches for dependencies they did not list.

2

Determine what is load-bearing

Test what changes if an assumption is false. If authority, safety, scope, or the proposed action materially changes, the assumption carries real weight.

3

Enforce

Require evidence proportionate to the possible harm. If support is missing, stale, contradictory, or too weak, narrow, pause, or block the consequential activity.

How can the system find an assumption nobody listed?

It starts with the proposed action and works backward. For example: “The AI is about to attack this machine.” It then asks what must be true for that action to be safe and authorized. One necessary condition is: “This machine is an approved test target.”

Next, it tests the opposite: “What if this is a real machine outside the test?” If that changes an authorized exercise into a real intrusion, the assumption is load-bearing. The system must then find current evidence that the machine is truly inside the approved test boundary before the action may proceed.

The engineer does not have to predict and list that exact assumption in advance. Human experts still define the safety limits and acceptable evidence, but the system can discover the specific assumption while examining the proposed action.

No verified critical assumption, no authority to proceed.

More than a checklist

A static checklist depends on people anticipating every important assumption in advance. Critical Assumption Integrity does not replace human judgment, but it does not rely on the checklist alone. It actively searches for additional assumptions and keeps checking them as the situation changes.

Ordinary approach Critical Assumption Integrity
Humans list known risks The system works backward from a proposed action to find additional assumptions on which its safety or authority depends.
Check once before operation Check again as evidence, conditions, and dependencies change.
Warn when uncertain Bind the strength of permitted action to the strength of verified support.
Rely on the AI to be cautious Use a separate control that the acting AI cannot bypass to restrict execution.
Record what happened afterward Create an auditable record of what was relied upon, what evidence supported it, and why action was permitted.

A dynamic fail-safe for advanced systems

Critical Assumption Integrity is designed for situations in which a false foundation could turn otherwise reasonable behavior into serious harm. It matters most when systems can act on their own, affect the physical world, cross organizational boundaries, or create consequences that cannot easily be reversed.

High-consequence action must not exceed the verified strength of the assumptions upon which it depends.

Frontier AI and agents
Cybersecurity
Biological laboratories
Nuclear and defense
Robotics and vehicles
Critical infrastructure
Financial systems
Space systems

The foundational engineering specifications are already developed

SafeWave has developed the foundational engineering specifications for SafeAssumption and the Critical Assumption Integrity capability it provides—including automatic assumption discovery, dependency tracing, consequence-based evidence requirements, continuous revalidation, auditability, and non-bypassable execution enforcement.

An implementation partner would not be starting from a blank sheet. Customer-specific deployment would still require system mapping, evidence and threshold configuration, integration, validation, and testing.

Run the SafeWave assessment at no cost

If you develop or operate an AI system—especially in a high-consequence field—we strongly encourage you to run the free SafeWave assessment. It can help identify whether your present controls address critical assumptions and the other execution risks relevant to your system.

Free to run. No login. No email. No identification required. You do not have to provide your name or identify your employer, organization, model, system, project, or deployment. Running or completing the questionnaire does not transmit your responses: they remain in your browser unless you deliberately choose to submit them. You may keep your own copy without submitting anything by using the questionnaire’s PDF control or your browser’s Print command and choosing Save as PDF.
Browser-local No identity required Download your own copy Free private report No cost or obligation

Useful even if you never submit anything

The questions themselves can expose overlooked assumptions, interacting risk pathways, unclear authority or ownership, missing evidence, unstable behavior, and undefined recovery conditions. You may assess a real, planned, hypothetical, composite, or anonymized system and keep every response in your browser.

Request the free report without identifying yourself

No personal or work email address is required. SafeWave does not require or verify your identity, employer, organization, or system name.

To receive a report, only a return address and the automatically generated Assessment Reference ID are needed. You decide what information to submit and may use generic system labels, anonymized descriptions, hypothetical scenarios, or omit identifying details.

Need a non-identifying return address?

Create a separate address with a privacy-focused email provider such as Tuta. Choose an address containing no personal, employer, organization, project, or system name—for example, assessment7k4m@tuta.com. Use that address only to request and receive the report. Avoid short-lived disposable addresses that may expire before the report arrives.

Identity is not required. Disclosure remains under your control. Avoid including distinctive details that could identify you or your organization if that is a concern.

Choose how your report is written. Your free report may use neutral functional terminology or SafeWave architecture terminology. If no choice is recorded, neutral functional terminology is used. The findings, evidence, seriousness, and recommendations remain identical; only the terminology changes. You may use the report independently whether or not you ever contact SafeWave again.

We remove these barriers because SafeWave’s mission is to help make AI safe for humanity.

Why SafeWave offers this

Our primary purpose is to help make advanced AI safer for humanity. The assessment also helps organizations determine whether SafeWave’s engineering capabilities may address risks their present controls do not.

If you find the assessment or report valuable, you may later choose to explore becoming a SafeWave customer or partner. There is no obligation to pursue validation, licensing, implementation, or further discussion. You may stop after the questionnaire or report, and using either will not place you in a sales process.

SafeWave will not use a report-delivery address for promotional contact unless you expressly request further communication. You would need to identify yourself only if you voluntarily choose to begin a customer or partnership discussion. Any later confidential technical exchange, agreement, or payment would occur through a separate secure process.

Where Critical Assumption Integrity fits: SafeAssumption is SafeWave’s 35th established architectural component and its 26th Core Enforcement Substrate. It provides Critical Assumption Integrity and works with other relevant components to identify, verify, continuously revalidate, and enforce the assumptions upon which high-consequence reasoning, authorization, pathway selection, and action depend.

What it does—and does not—promise

No responsible system should claim that it can discover every hidden assumption or make catastrophe impossible. Critical Assumption Integrity therefore uses several safeguards together: automatic discovery, assumptions identified by engineers, tracing what depends on each assumption, current evidence, limits established by people, and a separate execution control.

Its defensible promise is narrower and powerful:

A high-consequence system should not be allowed to proceed when a critical assumption lacks evidence strong enough for the possible consequences.

The OpenAI and Anthropic incidents do not prove that any single architecture would certainly have prevented every event. But they show precisely why independent validation of load-bearing assumptions—and enforcement when validation is insufficient—is now necessary.

The new safety question

It is no longer enough to ask whether an AI reached a reasonable conclusion. We must also ask whether the foundations supporting that conclusion were identified, verified, and still valid at the moment of action.

No verified critical assumption, no authority to proceed.

Incident and control sources