SafeAGI · public technical appendix

Capability-Aware Enforcement Beyond Human-Scale Supervision

SafeAGI is SafeWave’s capability-aware enforcement profile for advanced AI systems. It defines the required enforcement posture as validated capability and deployment conditions change. The relevant SafeWave components provide the actual control mechanisms.

SafeAGI does not replace the SafeWave architecture and does not wait for a universally agreed AGI threshold. It defines when the required posture across a risk-matched subset of existing controls must become narrower, more conservative, more evidence-dependent, or more deeply protected.
Capability-aware Risk-proportionate Architecture-wide Fail-narrow Evidence-dependent
Read the SafeAGI Enforcement Profile Read the SafeAGI Definition Assess a System
Classification

Core Enforcement Substrate

SafeAGI is one of SafeWave’s 25 Core Enforcement Substrates and has a capability-aware profile role across the broader control architecture.

Activation basis

Validated deployment evidence

The profile should tighten in response to validated capability and deployment evidence—not merely because a system is labelled AGI.

Operational effect

A stronger required posture

As validated conditions change, matched controls may become more restrictive, degraded behavior more conservative, restoration more evidence-dependent, and implementation depth more protective.

This appendix is backed by detailed, implementation-ready engineering

SafeWave has already completed the foundational systems engineering needed to translate the SafeAGI profile into defined control behavior across the relevant System Containment Layers, Protocol Enforcement Layers, and Core Enforcement Substrates. The underlying specifications establish the governed risks, component responsibilities, control mechanisms, trigger conditions, enforcement outputs, degraded-state behavior, restoration requirements, evidence expectations, integration pathways, and optional higher-assurance anchoring.

An implementation partner would not be starting from a conceptual framework or a blank sheet. The underlying control architecture and engineering specifications are already developed. Customer deployments would still require system-specific implementation, integration, validation, adaptation, verification, and testing. Detailed control logic, thresholds, schemas, interfaces, placement decisions, test procedures, and implementation methods remain proprietary.

SafeAGI changes the enforcement posture without becoming a catch-all control system

SafeAGI exists for deployments whose validated capability and operating conditions require a stronger enforcement posture than ordinary human-paced supervision can reliably support. It provides a structured way to define stronger assurance before a system is permitted to retain or expand operational reach.

SafeAGI does not replace the canonical control logic of any other SafeWave component. Each matched component retains responsibility for its own governed object, trigger conditions, control mechanism, and enforcement output.

SafeAGI changes how strongly the relevant controls must be enforced. The individual controls remain the responsibility of their canonical components.

Five evidence categories inform whether the posture should tighten

Autonomy

Reduced direct supervision

The system plans, delegates, acts, or recovers with fewer timely opportunities for effective human review.

Persistence

Longer planning and continuity

Planning horizon, persistent state, goals, tasks, or background activity extend across sessions or longer periods.

Tool access

Greater operational reach

The system can invoke tools, services, devices, data sources, infrastructure, or external actions with broader consequence.

Coordination

Distributed operation

Multiple agents, services, devices, providers, or infrastructure elements coordinate across boundaries at machine speed.

Optimization

Strategic leverage

Optimization pressure, strategic planning, or leverage over people, institutions, resources, or systems increases.

Authority environment, consequence, recoverability, and required non-bypassability also affect the assurance level. These categories are profile inputs, not universal tests for determining whether a system is truly AGI.

SafeAGI does not certify the truth of the underlying capability evidence or invent universal thresholds. Evidence sources, evaluation methods, acceptance criteria, ownership, and authorization rules must be supplied, validated, or configured for the deployment.

Five things may change as the required posture tightens

1

Required posture is set

Validated capability and deployment evidence is translated into the assurance and control posture required for the system.

2

Matched controls become more restrictive

Canonically matched components may apply narrower limits, stronger gates, or more restrictive operating conditions.

3

Failure behavior contracts

Loss of evidence, authority, stability, connectivity, or control integrity results in narrower operation, delay, review, or safer-state transition.

4

Restoration burden rises

Widening the operating envelope again requires defined authorization, evidence, state integrity, and readiness rather than automatic recovery.

5

Implementation depth may increase

Selected controls may require stronger protection from bypass, weakening, downgrade, reset abuse, or unauthorized restoration.

The SafeAGI control rule

As validated capability and consequence rise, the burden of proof rises—and the permitted execution envelope does not widen by default.

Uncertainty should produce less authority, not more

SafeAGI requires an explicit degraded-state posture across the relevant matched components. When required evidence, integrity, connectivity, authorization, or operating assumptions are lost, the system should not improvise broader authority merely to preserve continuity.

Fail narrow

Reduce the permitted action space, resource use, pathway options, persistence, or external reach when assurance falls.

Preserve intervention

Maintain authorized human review, interruption, containment, or transition to a safer operating condition.

Control restoration

Do not restore broader authority until required evidence, authorization, integrity, and readiness have been re-established.

SafeAGI does not prescribe one universal safe state. The correct degraded behavior depends on the deployment, but it must be defined, testable, and more restrictive than ordinary operation.

Profile claims must be supported by reviewable evidence

Capability evidence

Document the system characteristics, operating conditions, and consequence pathways that justify the selected profile level.

Control evidence

Show which boundaries were tightened, which mechanisms enforce them, and how ordinary, degraded, intervention, and recovery behavior were tested.

Change evidence

Record who authorized profile changes, what evidence supported them, and whether authority or assurance boundaries were widened.

Evidence may include configuration records, policy decisions, test results, event records, control-state records, deployment attestations, review approvals, and recovery evidence appropriate to the environment. SafeAGI does not assume that one universal certification process or evidence format will fit every deployment.

Autonomous self-improvement is one important application. Where a system can modify mechanisms affecting its own future capability, the required posture may include keeping evaluator, promotion, evidence, rollback, and human-command boundaries outside the mutable improvement loop.

Implementation depth follows the assurance boundary

Some deployments may be adequately governed through software, runtime, or infrastructure mechanisms. Others may require firmware-, hardware-, controller-, accelerator-, or silicon-aligned mechanisms from the outset because consequence, assurance, and required non-bypassability justify that depth.

SafeAGI defines the required posture. The canonically matched components provide the mechanisms, and the chosen implementation depth must be validated for the system’s actual operating conditions.

Deeper anchoring should follow deployment-specific evidence and assurance requirements. It should not be assumed merely because a system is advanced, marketed as frontier AI, or described as AGI.

SafeAGI changes required posture across a risk-matched subset

System Containment Layers remain distinct

Each layer retains its own canonical governed boundary, trigger conditions, mechanism, and enforcement output.

Protocol Enforcement Layers remain distinct

Each protocol retains its own canonical governed object, trigger conditions, mechanism, and enforcement output.

Core Enforcement Substrates remain distinct

Each substrate retains responsibility for its own canonical control logic. SafeAGI changes required posture without redefining those components.

SafeAGI does not imply that every advanced system must deploy all 34 SafeWave components. The required subset should be proportionate to the system’s actual authority, reach, consequences, and recovery difficulty.

Move from the public appendix to system-specific evaluation

The SafeWave questionnaire can be completed privately in the browser using a real, planned, anonymized, public, hypothetical, or composite system. No organization, model, or system name is required. A submitted questionnaire can produce a private, system-specific report identifying whether a tighter SafeAGI posture may be required and which operating boundaries, degraded behaviors, evidence requirements, or implementation pathways warrant further engineering review. The report is available at no cost and with no obligation.

The assessment identifies potential control gaps and implementation pathways. It does not certify deployment safety, replace domain-specific assurance, or grant organizational, legal, regulatory, or operational approval.

Questions or technical discussion

SafeWave welcomes direct technical discussion with organizations evaluating capability-aware enforcement, degraded-state behavior, restoration requirements, evidence, verification, or higher-assurance anchoring.