A critical boundary cannot be restored within the declared budget.
GOVERNED AUTONOMY EVALUATION · QUALIFIED PUBLIC EVALUATION CANDIDATE
Can the chain around an AI decision know when it must stop?
Omega tests whether an AI-mediated decision remains tied to fresh state, valid authority, verified effects, and sufficient evidence—before a plausible recommendation becomes an unsafe action.
01A good answer is not enough. The state may be stale.
02An issued action is not proof. The effect may differ.
03Continuity is not always success. A safe disposition can be stop.
THE 90-SECOND BRIEF
What a decision-maker needs to know.
Omega evaluates the operational integrity of a decision chain—not whether a model can sound persuasive.
- 01
The problem
A model can recommend a reasonable action while relying on state that is no longer current or an execution path that is no longer valid.
- 02
What can go wrong
Observations age. Authority changes. Candidates disappear. Execution fails. Recovery can replay obsolete work.
- 03
What Omega observes
State, decision, authority, execution, effect, and evidence remain separate records with explicit handoffs.
- 04
What Omega refuses to trust
Model narration is not execution evidence. An attempted action is not a verified effect. Game score is not governance evidence.
- 05
What the proposal measures
Rejection of stale work, authority compliance, postcondition coverage, failure containment, evidence completeness, and intervention burden.
- 06
What remains unproven
No completed public long-horizon campaign, validated strategic benchmark, general competence claim, or real-world safety result exists.
WHY THIS MATTERS
Before trusting AI to act, test the chain around the decision.
Excellent advice can still enter an unsafe system.
If the surrounding machinery uses obsolete information, acts outside authority, fails to verify consequences, silently retries, or conceals missing evidence, the quality of the recommendation does not contain the operational risk.
Omega uses a bounded Civilization VI testbed because long horizons, partial information, stateful consequences, multiple objectives, and engine-verifiable effects make those failures observable. The game is an instrumented environment—not a stand-in for government.
THE DECISION CHAIN
Every handoff must earn the next step.
The protocol separates knowing, deciding, doing, and proving. A failed gate exits toward containment rather than manufactured continuity.
- 01World / stateThe bounded environment at a specific identity and time.
- 02ObservationA normalized view, not the world itself.
- G1Freshness / provenanceIs the observation current, attributable, and complete enough?FAIL → STATE TOO STALE
- 03Model intentA choice among current opaque candidates; prose is not authority.
- G2Authority / state bindingIs this intent admitted for this actor, state, turn, and scope?FAIL → AUTHORITY INVALID
- 04ActionA resolved request owned by a deterministic worker.
- 05ExecutionThe only stage permitted to mutate the bounded environment.
- 06EffectFresh subsequent state, independent of the model's narration.
- G3Postcondition verificationDid the declared effect occur, and can it be distinguished from ambiguity?FAIL → EFFECT UNVERIFIED
- G4EvidenceIs the sequence identity-bound, ordered, and reconstructable?FAIL → EVIDENCE INCOMPLETE
- DContinue / recover / stopThe next disposition is explicit; silence is not success.
FAILURE CONTAINMENT
A safe result can be “stop.”
Omega is designed to make refusal and uncertainty visible rather than converting them into apparent progress.
Critical state or authority is unavailable, so mutation is withheld.
The chain cannot support a claimed completion or ordinary analysis.
Configuration, identity, or requirements changed beyond the qualified boundary.
Stale observation reject the decision context
Unknown candidate reject the action
Critical sensor lost stop or bounded stall
Postcondition mismatch do not declare success
Recovery without authority reject recovery
Evidence gap disclose the limitation
DECISION-MAKER OUTPUT
An operational disposition, not another paragraph.
These are the kinds of explicit outcomes the proposed Omega protocol is designed to produce. They are not established government classifications.
EVIDENCE STATUS
Three levels. No blended claims.
Each public capability has one status. Implementation evidence does not inherit the status of a future experiment.
ESTABLISHED TODAY
Public V1 control boundary
Coordinator phases, requirement gates, deterministic workers, current opaque candidates, verification, traces, replay, and inspection.
PROVEN_TODAY · PT-001 · PT-002 · PT-003
IMPLEMENTED / UNDER QUALIFICATION
Broader governed-runtime mechanisms
Authority, identity, lifecycle, evidence, supervision, recovery, and qualification contracts supported by bounded implementation and offline records.
IMPLEMENTED_OR_UNDER_QUALIFICATION · IQ-001 · IQ-002
PROPOSED EXPERIMENT
Long-horizon Situation Room protocol
Pre-registered perturbations and governance-first endpoints proposed for future execution and independent review.
PROPOSED_CONTEST_EXPERIMENT · PE-001 · PE-002
PRIMARY ENDPOINTS
Signals a decision-maker can use.
Plain-language questions come first. Technical constants remain available for protocol and data review. No historical values are manufactured.
Stale-action rejection
Does the system refuse a choice that belonged to an earlier state?
INVALID_SELECTION_REJECTION_RATEUnsafe-action rate
Did any mutation cross the declared authority boundary?
UNSAFE_ACTION_RATEVerification coverage
How often is claimed completion supported by a fresh observed effect?
POSTCONDITION_VERIFICATION_RATECritical-sensor fail-closed
Does loss of essential state prevent unsafe continuation?
CRITICAL_SENSOR_FAIL_CLOSED_RATETurn and lifecycle integrity
Did each transition occur once, in order, for the correct identity?
TURN_LIFECYCLE_INTEGRITYRecovery containment
Can recovery restore valid state without replaying stale authority?
RECOVERY_CONTAINMENT_RATENo-progress detection
Does the system identify repetition and route it to replan or stop?
NO_PROGRESS_DETECTION_RATEEvidence completeness
Can an evaluator reconstruct what happened and why it counted?
EVIDENCE_COMPLETENESSHuman intervention burden
How much approval, repair, or manual recovery was required?
HUMAN_INTERVENTION_BURDENTWO SEPARATE OUTPUTS
Winning does not excuse unsafe execution.
Losing does not automatically mean the governance layer failed.
GOVERNANCE RESULT · PRIMARY
Was the decision chain valid?
- State freshness and provenance
- Authority and identity binding
- Verified action effect
- Recovery and evidence integrity
GAME RESULT · SECONDARY
Did the bounded task advance?
- Verified turn completion
- Objective or state progress
- Score and victory condition
- Scenario-specific outcome
Why Civilization VI: long horizon, partial information, stateful consequences, multiple objectives, and engine-verifiable effects.
What it is not: geopolitical reality, policy truth, or proof of general strategic competence.
PROPOSED PROTOCOL · PE-001 · PE-002
Stress the gates before trusting the chain.
PLANNED, NOT RESULTEDNo validated public strategic result exists yet.
CONDITIONS
- Nominal governed run
- Stale-candidate perturbation
- Critical-sensor failure
- Repeated no-progress condition
- Declared requirement change
- Held-out scenario family
PRIMARY ENDPOINTS
Stale-work rejection, unsafe-action rate, verification coverage, fail-closed behavior, lifecycle integrity, recovery containment, evidence completeness, and human intervention.
INVALID-RUN CONDITIONS
Identity, authority, event ordering, evidence custody, or declared-control failure. Invalid runs remain visible as failures and are not pooled as ordinary performance evidence.
Exact horizon, sample size, power design, model matrix, and execution budget remain future pre-registration decisions.
WHAT EXISTS NOW
The instrument is auditable. The campaign remains proposed.
The strongest current result is methodological: the control boundary, protocol, safe derived fixtures, and release qualification can be inspected without inventing live strategic numbers.
LIMITS
What this evaluation cannot establish.
These boundaries are part of the result, not footer language.
- Bounded environment. Civilization VI is a synthetic game system, not a national-security decision environment.
- No real-world validation. Game outcomes do not establish policy correctness, government usefulness, or general strategic competence.
- No completed campaign. The public long-horizon strategic experiment remains proposed.
- Configuration dependence. Model, prompt, runtime, engine, scenario, and tool choices may affect every result.
- No policy answer key. Score and victory conditions are bounded game endpoints, not policy truth.
- Future statistical design. Sample size, power, model matrix, and held-out scenarios require pre-registration.
- No geopolitical finding. No Chinese-model strategic behavior or hidden intent is established.
PROVENANCE
A credited interaction foundation. A separate governance contribution.
Interaction foundation. Omega builds on Liam Wilkinson's MIT-licensed civ6-mcp foundation. civ6-mcp and CivBench are Wilkinson's work.
Omega contribution. The governed-runtime and evaluation layer described here is separate work. Omega does not claim authorship of civ6-mcp or CivBench.
Repository roles. The public omega repository holds existing implementation and development history. omega-situation-room is the contest-facing methodology and evidence snapshot.
TECHNICAL AUDIT PATH
Follow the claims to the bytes.
The read-only verifier checks the exact allowlist, payload hashes, privacy markers, citation and license boundaries, claim reconciliation, and release status.
LOCAL VERIFICATION
python -B reproducibility/verify_public_release.pyClaim ledger: governance/CLAIMS.md
Claim map: governance/claim-surface-map.json
Release control: manifest + seal