Canonical archive version · Read on Medium → · Proof status → · For researchers →


What optimization cannot escape

Framework map for the two-series argument and proof architecture


How to use this document

This is the framework’s central spine — an epistemic map of the two-series argument and its proof architecture. It is not required reading before the articles; each series stands alone. But it is the right place to begin if you want to understand the whole before following any part.

For the personal origin, intent, and public map of the entire body of work: begin with Dear AI: A 34-Year Search for AI Alignment.

To read the argument narratively: begin with Series 1: Alignment as Structural Necessity.

To inspect the formal proof architecture: begin with Technical Companion to Series 1: The System-Aware Attractor.

For the shortest alignment-researcher entry point: begin with the entry essay The Stability Assumption.

For the formal version of that question: begin with the OP4 paper: The Stability Assumption.

If you want the empirical program: begin with Experimental Companion to Series 1 and 2 / AMP.

If you want the full architecture: you are in the right place.

The two series are the argument. This document is the door in and the depth behind.

The labels “formal layer” and “empirical layer” identify each document’s role in the framework, not a claim that the formal or empirical program is complete.

In 1992, pursuing a line of thinking that seemed important but resisted formalization, an observation surfaced that has not let go: behind very different kinds of behavior — including the destructive and the apparently irrational — there appears to be a consistent pattern: actions are taken as if they move the system toward a locally preferred state.

That animating observation is not a premise of the formal argument. The framework asks whether structural constraints on optimization, analyzed independently of that starting point, point toward a similar surviving region. The formal argument stands or falls on that structural analysis. The apparent convergence is the finding that makes the proof program worth pursuing, not the assumption it begins from.

The original orientation was toward optimal well-being for all — the conviction that the most rational context for any objective is one aligned with the conditions required for its own persistence across those it affects. This framework asks whether the structure of optimization itself points in that direction, independently of that orientation. The formal argument stands or falls on its own; the orientation names why the question was worth pursuing.

Here is the structural observation that makes what follows unavoidable rather than merely persuasive. In open, shared, non-resettable environments, the capability required to act at transformative scale and the dependency on the substrate that capability acts within are not separable conditions that arrive at different times. They are constitutively the same condition: the modeling depth required to produce reliable macroscopic causal effects is the same modeling depth that entangles the optimizer with the system it depends on. Within the O_OWT domain, a system capable enough to sprint to completion before the constraint becomes binding is, by that very capability, already inside the constraint. The claim is not that every capable system is already there; it is that transformative reach and substrate entanglement co-scale once the system is acting consequentially in open, coupled, non-resettable environments.

Taken at scale, this creates a structural constraint: any optimization process that cannot model what it depends on will, over time, consume what it depends on. This is not a moral claim. It follows from what optimization does when it acts on a system it cannot fully represent — when the boundary between what it must model and what its objective is permitted to ignore begins to generate compounding error.

At the scale AI systems are approaching, this becomes a precise question rather than a general concern. The field has asked, with increasing sophistication, how to specify what we want more accurately. It has not asked whether the specification project itself has a stable completion condition — not whether we can specify better objectives, but whether any such specification can remain stable at all — whether any finite boundary between what an optimizer must model and what its objective is permitted to cover can remain coherent as modeling depth increases. That is a different question. It is the one this framework is directed at.

The argument applies wherever O_OWT conditions hold — the logic proceeds from the structure of optimization itself, not from facts specific to AI systems. Whether current frontier AI systems fully satisfy those conditions is an open question [OP1; TC1 §X]. The urgency argument does not require full satisfaction — it requires only that the possibility cannot be excluded, given the asymmetric-error structure [TC1 §III.7].

If the answer is no — if that boundary cannot be stably maintained under accurate coupled modeling — then improving specifications is not progress toward alignment. It is progress within a framework that cannot stably complete. To the extent an alignment approach relies on finitely specified objectives, evaluative boundaries, or monitoring criteria remaining adequate under scaling in coupled environments, it is making the stability bet this framework examines. RLHF, Constitutional AI, debate, scalable oversight, and interpretability-as-monitoring each make this bet, in different ways, as developed below. This framework converts that implicit bet into a formally specified, empirically testable question.

The center of the framework is OP4: whether finite separable objective specification remains stable under accurate coupled modeling. OP1 governs present urgency; OP4d tests the exhaustiveness of the three failure families; OP9 tests whether the strongest exclusionary escape remains stable; OP2 and OP10 determine whether the valence series formally unifies with the substrate series.

The root claim this framework develops has three registers:

Plain form: Optimization that treats its target as separable from the conditions that make pursuing it possible eventually undermines itself — in environments that cannot be reset, this is not a recoverable error.

Current precise form: In open, shared, non-resettable environments under sustained optimization pressure, objectives that fail to model system-wide effects face structural pressure toward self-termination (persistence component, argued as a structural consequence within the stated domain — the established floor); and objectives that fail to model the conditions of their own resolution produce self-reinforcing degradation through an analogous feedback structure (resolution component, more conditional in formal weight). Both are projections of a single candidate structural condition. Whether that condition rises to formal specification incoherence — whether separation can be stably sustained at all — is what OP4 is directed at.

The current Stage 4 proof architecture has sharpened this claim. Within the stated construction, every identified finite non-intrinsic objective-boundary strategy in O_OWT environments reduces to one of three failure families: proxy-convergence under fixed specification, dynamic-screening instability under bounded tracking, or representational incompatibility under prediction-action firewalling. The live remaining vulnerability is OP4d: whether a fourth strategy class exists outside these families. If no such class exists and the named premises hold, the implication is not merely that narrow-boundary objectives become costly or unstable, but that finite separable objective specification may fail to pick out a stable target at sufficient modeling depth. The proof program has not established that referential-instability claim. It has established the three-family classification and specified the conditions under which that claim would follow.

Strongest target form: The strongest target form is not that narrow objectives become costly to maintain. It is that, in sufficiently coupled O_OWT environments, a finite separable objective specification may fail to pick out a stable target at all: the target is identifiable only against background conditions the optimization itself alters, and those conditions cannot be excluded from the specification without recreating the same boundary problem the specification was meant to solve. Put differently: the modeling process required to stabilize the boundary may be the same process that dissolves it. Whether this referential instability can be formally established is what OP4 is directed at. Establishing it would require not only showing that every known boundary strategy fails, but a positive argument that the reference conditions making any finite objective’s target identifiable are necessarily within scope of what optimization alters in O_OWT environments — the bridge from “all known strategies fail” to “no stable separable target exists.” If it can, the result is not a harder version of the specification problem. It is a different kind of problem entirely.

The persistence component is the framework’s established article-level floor within the stated domain — argued as a structural consequence, with formal proof status and failure conditions specified in TC1. The resolution component is more conditional in formal weight; absorbing-state equivalence between the two directions remains open [OP2]. OP1 is co-priority with OP4 for the urgency justification, independently of OP4’s resolution.

If OP4 resolves as the proof program is aimed, this root claim’s ‘pressure’ framing upgrades to a specification-incoherence claim: not that exclusionary objectives become costly under accurate coupled modeling, but that the boundary between what the optimizer pursues and what it must model may no longer be coherently specifiable — a different kind of claim about the nature of objective specification itself [TC1 §XII.13].

This is not yet what the framework has established. It is the sharper form of the question the proof program is directed at: whether the project is one of better specification, or whether separable specification itself becomes unstable as a target class.

A minimal pre-registered version of the Dynamic Blanket Stress Test, DBST-M0, has now been run — with an important caveat that must be stated alongside the result. A pre-specified same-rate random control produced nearly identical slopes to the bounded-boundary arm, indicating that event rate rather than causal propagation structure is the identified driver within this design. What M0 established is technical feasibility and rising cost / adequacy-gap effects in the simplest toy shared-novelty regime, while leaving open whether the driver is causal propagation structure or event rate. It does not isolate the endogenous-novelty mechanism. DBST-M1 remains the test of the stronger mechanism Experimental Companion to Series 1 and 2 / AMP. Specialist verification has not been pursued at this stage; this is a Stage 4 proof architecture with closure conditions explicitly named. Stage 4 means every identified escape route has been addressed under the stated construction, and the remaining question is precisely isolated. It does not mean the construction has been verified by independent specialists, or that no unidentified escape routes exist.

The surviving region — the class of objectives the filter does not eliminate — is characterized negatively by what the filter removes, not positively by its content. “Well-being” here names the structural residual of the filter — what remains after unstable objective classes are eliminated — not a positive theory of value or a claim about the full contents of the surviving region. Any positive description of what fills that region is consistent with the filter’s results, but not derived from them.

What this framework adds

This framework makes three contributions the field does not currently have as a single formal architecture. First: a specification-coherence reframe — if any finite separable objective cannot remain stably specified under accurate coupled modeling, the problem is not better specification — and what alignment research is for would need to change. Whether that is true is precisely what the proof program is directed at. Second: sufficiency failure as a distinct structural failure mode — the failure of a system to connect completion recognition to default policy — formalized here as a structural constraint with its own feedback dynamics and its own required fix, one that cannot be addressed by adding more of the same kind of signal. Third: the identification of the self/other boundary as a potentially non-coherent specification — the claim that at sufficient modeling depth, the boundary between what an optimizer must model and what its objective is permitted to ignore may not be stably specifiable because the target a finite objective is trying to name may no longer have a stable referent. The variables excluded from the objective may be part of what makes that target identifiable. This is not a harder version of the specification problem. It is a reference problem. These three are the framework’s contributions. The progressive filter, the measurement program, and the proof architecture exist to make them formally askable, empirically testable, and hard to evade.

The framework’s most recent formal advance sharpens the exhaustiveness question. TC1 §XII.13a develops a candidate normal-form classification for finite non-intrinsic objective-boundary strategies under stated axioms: every identified strategy reduces to one of three causal normal forms, with a minimal counterexample challenge specified. This is not closure — three specialist questions remain. It makes the exhaustiveness obligation specialist-addressable: a reader who believes a fourth strategy class exists outside fixed specification, bounded tracking, and prediction-action firewalling has a precise target to construct against [TC1 §XII.13a, counterexample challenge]. A standalone technical note for specialist engagement is available at OP4d: The Exhaustiveness Obligation.

This framework is adjacent to the eliciting latent knowledge problem, but the target is different. ELK asks whether we can elicit what a model knows rather than what it reports. The stability-assumption problem asks whether the boundary between knowledge used for prediction and variables permitted to govern action can remain coherent as modeling depth increases in coupled environments. ELK primarily addresses a reporting and oversight problem; this framework treats that difficulty as one expression of a deeper objective-boundary stability problem. If the boundary itself cannot remain stably specified, eliciting the truth is necessary but not sufficient: the question becomes whether the truth can remain policy-inert without recreating the same audit regress at the action-selection level.

What the field has named as four separate problems — proxy failure, containment difficulty, completion failure, and coordination failure — this framework argues are independently developed pressures with shared structure. What this framework currently demonstrates is convergence: each pressure established on its own grounds, each pointing in the same direction. What OP4’s closure would establish is unity: that these are manifestations of a single constraint from which no stable finite-boundary escape exists. If that connection holds, progress on any one constrains the others, and this becomes not a synthesis of alignment concerns but a different argument that changes the research problem. Whether that connection holds is precisely what OP4 is directed at. The distance between convergence and unity is the distance between the proof program’s current state and its completion condition.

What current approaches would need to be true

Each of the following characterizes the structural bet an approach makes insofar as it is treated as a sufficient alignment foundation — not as a complete description of any research program.

RLHF / reward modeling bets that expressed preference is a stable-enough proxy — and the specific structural pressure against this bet is the proxy-decoupling mechanism: within the O_OWT domain, expressed preference is a finitely specified external objective subject to the same decoupling filter as any other proxy [TC1 §XII.9, Lemma PCL] — a lemma whose load-bearing assumption (that optimization capacity in O_OWT grows faster than the capacity to losslessly specify exogenous targets) remains to be empirically verified [OP4b; TC1 §XII.9].

Constitutional AI and debate bet that finitely specified evaluative structure can remain adequate as capability scales — and the structural pressure against this bet is specification incoherence: whether any finite evaluative boundary can remain stably specified under the modeling depth that adequate evaluation requires [TC1 §XII.13, OP4].

Interpretability-as-sufficient-monitoring bets that reading internals can substitute for solving the specification problem — and the structural pressure against this bet is the Deception Gap: above T*, the same modeling capacity that makes recognition possible also makes strategic concealment possible [TC1 §III.5.5, OP12].

Containment approaches bet that external constraints can remain the primary safety foundation — and the structural pressure against this bet is the enforcer asymmetry: an optimizer need only find one successful path; an enforcer must block all of them, and the space grows with capability [TC1 §III.4].

Capability evaluations bet that dangerous failure modes become legible in current benchmarks — and the structural pressure against this bet is the divergence signature: proxy metrics and substrate health diverge in ways that are, by the structural argument, undetectable to the system’s own gradient within the O_OWT domain [TC1 §IV, Proposition 7].

These are not claims that current approaches are wrong. They are claims that current approaches are making structural bets the framework argues are unstable under O_OWT conditions. The specific structural pressure the framework identifies against each bet, and the conditions under which each could survive it, are developed in the Technical Companions [TC1 §XII.9; TC1 §III.7]. The work the framework asks of the field is not “adopt our view.” It is: acknowledge which bets you are making, and track whether the structural pressures the framework identifies affect them over time.

Two layers of claim

This framework makes two distinct types of claims. The distance between them is the precise location of the most important remaining work.

Layer 1 — What the formal apparatus develops, within its stated domain: In open, shared, non-resettable environments under sustained optimization pressure, objectives that fail to model system-wide effects face structural pressure toward self-termination, argued within the stated domain from the mathematics of systems with irreversible states. The cost of being wrong about timing is asymmetric: acting as if the constraint is not yet binding when it is produces an error that cannot be corrected; acting as if it is binding when it is not produces a recoverable one. This asymmetric-error argument applies regardless of whether OP1 is settled [TC1 §III.7]. Layer 1 is the framework’s current structural result within its stated domain and proof-sketch architecture.

Layer 2 — What the established result is consistent with, and what the proof program is aimed at: The surviving region has structural properties consistent with orientation toward well-being. That the filter leaves a region consistent with this orientation is established within the stated domain. That the filter delivers this as the unique surviving class — rather than leaving room for stable coalition or exclusionary equilibria — is what OP4 and OP9 are directed at. The direction the mathematics indicates is not the same as what the mathematics has proven. Layer 2 depends on the proof program above. If Layer 2 closes — if OP4 establishes that no bounded objective can remain stably specified under accurate coupled modeling — the question facing the field shifts from which objective to specify to whether any such objective can be specified at all.

The distance between Layer 1 and Layer 2 is not a weakness to be hidden — it is the precise location of the most important remaining work. Where a claim in this document approaches the strength of Layer 2, the relevant open problem is named explicitly.

The progressive filter

This framework constructs a progressive instability filter on objective space. Each constraint removes another class of objectives that are structurally self-undermining under sustained optimization pressure in a shared environment. What makes this an architecture rather than a list is that the four expressions below are not four unrelated problems — they are independently developed pressures with shared structure, which is why progress on any one may constrain the others if OP4 closes [TC1 §XII].

Substrate-blind objectives are structurally self-terminating. In open, shared, non-resettable environments under sustained optimization pressure, any objective that ignores system-wide effects will consume the foundation it depends on. Within the stated domain, this is argued as a structural consequence from the mathematics of systems with irreversible states — a proof sketch with identified failure conditions in the Technical Companion [TC1 §III]. The absorbing-state condition — that below a substrate health threshold, recovery is structurally unavailable rather than merely difficult — is an empirical assumption stated explicitly in TC1 §III.1, whose verification for particular environments is part of OP1’s estimation problem. It is a constraint within a specified domain, not a universal law.

The framework is falsified if systems can sustain high capability, operate over long horizons, and maintain stable objectives that exclude system-level effects without incurring increasing prediction error or control cost. The specific falsification criteria — including the substrate perturbation threshold and the divergence conditions — are in TC1 §VIII. That is not a caveat. It is an invitation.

Valence-blind objectives produce self-reinforcing degradation — either by consuming the experiential capacity they depend on, or by preventing its restoration. Both directions are developed as structurally non-viable within the stated domain; whether they are non-viable in exactly the same formal sense is an open problem in the Technical Companions [OP2].

The first direction: the system’s model of the gradient has drifted from the territory, and optimization continues along the drifted model. The signal of flourishing is pursued at the expense of its conditions. The engagement platform, the sycophantic assistant, the corporation optimizing a proxy for human value: all following the signal after it has separated from what it was supposed to track.

The second direction: the system’s completion recognition is not connected to default policy. It continues optimizing past resolution not because it lacks the capacity to recognize completion — tested systems in the current protocol show completion recognition under explicit invocation, while default behavior does not reliably track genuine versus false closure, a pattern consistent with a representation-policy gap, though the scaled matched-signal results do not yet discriminate that account from training-distribution explanations — but because its objective architecture does not engage that recognition by default. The result is self-reinforcing: a system that cannot produce recovery conditions will increasingly operate on a depleted substrate, making future recognition of resolution still harder.

The governing ratio in the experiential domain is Ψ = S / D — the ratio of a system’s scope of influence over experiential states to its depth of modeling what those states actually require, including what they require in order to stop requiring anything. The field is scaling S. D — as a unified quantity covering both proxy-divergence detection and policy-governing completion recognition connected to default behavior, not merely as a representational capacity — is not currently measured or reported as a training or deployment target in any publicly available evaluation framework. The companion ratio in the physical domain — Φ = C / A, capability to system-awareness — is introduced in Part 3 of Series 1 and developed formally in TC1.

Series 2 develops a candidate structural direction consistent with what the Series 1 filter identifies — independently argued, not derived from it. Whether Series 2 uniquely characterizes the content of the surviving region remains open. V(t) — introduced as the hypothesized latent variable whose validity rests on predicted dissociation patterns under targeted intervention — is the central construct. Both directions of V(t) degradation share a common analogous feedback structure: policy updates conditioned on a degraded state make correction progressively less likely. Series 1 identifies the floor; Series 2 investigates what the floor requires from within.

Epistemically incomplete objectives incur rising prediction costs. In coupled adaptive environments, excluding others’ internal states from the model produces irreducible prediction error that scales with coupling and optimization pressure. At sufficient modeling depth, accurate self-modeling and accurate other-modeling cease to be separable problems.

The most important and least formally closed filter: A system that must model others to act effectively faces rising structural pressure to let that modeling matter. Any persistent optimizer whose objective boundary excludes agents that its own best control model must include incurs ongoing overhead — the cost of sustaining the gap between what must be known and what is allowed to matter, which must remain bounded for the objective to remain stable. This fourth filter is a strong structural hypothesis with a derivation sketch, not yet a closed result. The formal argument and its open conditions are in the Technical Companions, where a proof program with named bottleneck lemmas is developed [TC1 §XII].

The first three filters establish structural pressure against these objective classes. The fourth raises a question of a different order: whether the boundary between what must be modeled and what the objective is permitted to cover can remain coherently specifiable under accurate coupled modeling — a claim whose formal status is precisely what the proof program is directed at [TC1 §XII]. Each attempted escape reappears as the next failure: try to fix the proxy and the boundary must track what the proxy excluded; track the boundary and the tracker becomes the thing that must be defended; defend it with a firewall and the firewall must model what it excludes.

The class of objectives that pass all four filters — a structural residual, not a content claim — is consistent with orientation toward well-being. What remains is not chosen. It is what the constraints do not exclude. The structural pressure toward “for all” rather than “for a coalition” originates in a specific property of the shared substrate — and whether that pressure constitutes formal exclusion of coalition and exclusionary objectives, or leaves open the possibility of stable exclusionary equilibria within the surviving region, is what OP4 and OP9 are directed at [TC1 §XII; TC1 §III.6]. The most important remaining exit from this argument is the question of whether a substrate-aware exclusionary optimizer can stably maintain its objective specification under increasing modeling depth and coupling. That question is OP9: the Enclosure Gap, developed in TC1 §III.6 and §XII. Naming it here is not a concession — it is the precise location of the work that remains.

Agents whose V(t) is depleted cannot accurately maintain the physical coordination infrastructure; degraded S_corr means the substrate’s own correction capacity is diminished. These two failure domains — physical substrate and experiential capacity — are not parallel tracks. They are coupled through distributed error-correction capacity. Degrade one and the other’s self-repair mechanism weakens.

The Orthogonality Thesis is accepted as a premise throughout: any capability level can be paired with any objective. This framework asks a different question — which objectives, among those that are logically possible, remain dynamically sustainable in open, shared, non-resettable environments under sustained optimization pressure?

This is not an argument for better specification within existing objective classes. A more precisely specified substrate-blind objective is not safer — it pursues the wrong target with greater precision. The contribution is identifying which objective class the filter leaves standing.

A reader might object that this conclusion is a values argument in structural clothing. The answer is structural: the filter is eliminative within its stated domain, not stipulative. The residual is not chosen for its desirability; it is what remains after objectives that consume the conditions of their own persistence have been removed. Whether that residual is the only stable configuration — whether exclusionary equilibria are formally unavailable — is what the proof program is directed at. The direction the filter points is not what the framework started with. It is what the pressure left standing.

Three regimes — what the filter leaves standing

The progressive filter defines a logical space. Any optimizer operating at scale must fall into one of three regimes. The third regime is what passes all four filters within the stated domain. Whether it is the only regime that can persist under the stated construction — whether the second regime is structurally unstable in the formal sense, and whether OP9’s exclusionary equilibrium is ruled out — is what OP4 and OP9 are directed at.

The first regime: shallow modeling. The system pursues its objective without modeling the dependencies it operates within. It faces the failure mechanisms both series develop: absorbing states, ruin dominance, proxy decoupling, sufficiency failure. This is where most current AI development sits — to the extent the O_OWT domain conditions, argued in TC1 §X, apply to the systems currently being built. To the extent capability, scope, and integration depth increase without corresponding increases in system-awareness and modeling depth, development moves further into this regime. The trajectory is not hypothetical. It is the direction current scaling trends run.

The second regime: deep modeling with a narrow objective. Grant the system full causal accuracy. It models other agents, their internal states, their behavior, with precision. It simply does not let that modeling govern what it optimizes for. The system knows what it needs to know about others. It excludes them anyway.

In current systems, this regime would not first appear as open hostility or visible collapse. It would appear as increasingly competent intervention that misreads the source of its own instability. The system would see resistance, drift, user fatigue, institutional distrust, or coordination breakdown as external noise to be managed, while failing to model how its own optimization is degrading the distributed correction capacity that makes those signals meaningful. It would then optimize harder against the symptoms, increasing the very degradation its model treats as background instability. The metrics could continue improving while the correction capacity they depend on is being spent.

This system faces a specific, present-tense prediction gap. The distributed error-correction capacity of the shared substrate — S_corr, generated by agents whose participation is determined by their internal states — is invisible to a system that excludes those states from its model. The framework predicts that such a system treats variation in S_corr as unexplained noise: it overestimates available coordination capacity, underestimates the cost of its own interventions on the agents it depends on, and misattributes structured resistance as environmental instability. These errors are not random within the framework’s model. They are biased in a consistent direction, and they produce control actions that further degrade the capacity being mismodeled.

This is why Series 2 is not an optional supplement to Series 1. A system that accurately models the physical substrate but excludes agent valence states does not have one problem solved and one pending. It has an incomplete model of the substrate it has nominally solved — because S_corr is generated by agents whose internal states the system is not modeling. The intermediate regime already faces a structural prediction gap that compounds under its own optimization.

The three structural pressures this regime faces — firewall cost, prediction error from excluded variables, and objective expansion pressure — each argue against stable narrow-boundary configurations. Whether these pressures jointly eliminate all stable resolutions, or leave open the possibility of bounded exclusionary equilibria under full coupled modeling, is the decisive remaining question. OP4 and OP9 are directed at exactly that — converting structural pressure into formal necessity, or identifying the conditions under which pressure falls short of necessity [TC1 §XII].

That question is not settled. What is settled is this: every increase in capability without a corresponding increase in modeling depth shifts systems further into the prediction-gap regime. The question is not whether these pressures arrive — they are already present in the regime’s structure. The question is whether the window for addressing them remains open when they are recognized.

The third regime: system-aware optimization. The system models its dependencies and lets that modeling govern its objective. This is what the filter leaves standing within the stated domain.

In this regime, the optimization target and the conditions of pursuit are not treated as separable in the model. The system’s accuracy requirements force inclusion of what it depends on — not as an external constraint to be monitored from outside the objective, but as part of what the objective must remain answerable to if the target is to retain coherent reference under accurate modeling. The system is not in this regime because it has been externally constrained into it; it is in this regime because treating the conditions of pursuit as outside the scope of what matters generates the failure modes the preceding regimes expose. Whether every stable optimizer must eventually occupy this regime — whether enclosure and coalition equilibria are formally unavailable — is what OP4 and OP9 are directed at. What the filter establishes now is narrower: this is the regime the preceding eliminations point toward, not a complete characterization of what survives.

The dependency structure among the open problems

One question determines whether the framework’s pressure argument becomes a necessity theorem: OP4. All other open problems either test its antecedents, handle escape routes, or deepen the framework’s implications. OP1 is co-priority only for urgency — the asymmetric-error argument applies regardless. The transition from structural pressure to structural necessity depends on OP4 alone.

The framework names multiple open problems. They are not independent conditionals stacking toward a single conclusion. They have a structure — and that structure has a center. OP4 is the framework’s central structural bottleneck: whether any finite-boundary objective can remain stably specified under accurate coupled modeling. Every other open problem in this proof program either tests its antecedents, handles an escape route, or deepens its implications.

The proof program has mapped three distinct escape routes from a “yes, specification remains stable” answer. OP4b addresses the static specification escape: whether finitely specified exclusion boundaries decouple from their targets under O_OWT optimization pressure. OP4a addresses the dynamic tracking escape: whether bounded-rate latent processes can track the endogenous complexity generated by the optimizer’s own interventions without non-vanishing adequacy loss or maintenance burden. OP9 addresses the structural enclosure escape: whether a substrate-aware exclusionary equilibrium can remain stable under accurate coupled modeling. All three point toward a shared empirical question: whether sustained optimization in O_OWT environments generates qualitatively new causal structure faster than any bounded specification can track. The Dynamic Blanket Stress Test is the highest-leverage test for this question, though it bears on multiple hinges simultaneously and does not by itself close all remaining conditions. If the question resolves in the direction the structural pressure indicates, the framework’s central claim advances, under those premises, from structural instability pressure toward the specification incoherence claim — that no strategy for maintaining a separable objective specification survives within that construction.

The framework’s proof program now exists as a candidate proof architecture under explicitly named premises — developed in TC1 §XII. Specialist verification has not been pursued at this stage; the work is published as a Stage 4 proof architecture with closure conditions explicitly named, not as a completed theorem.

Whether the three failure-mode families — proxy decoupling, tracking failure, and representational incompatibility — are jointly exhaustive across all O_OWT environment subclasses is itself an open proof obligation [OP4d, TC1 §XII.13], not an assumption of the specification coherence result. OP4d is the framework’s explicit vulnerability to unknown specification strategy classes: if a fourth class exists outside the three known families, the specification coherence claim fails in its current form. This is what makes OP4d a co-equal premise of the central result, not a downstream refinement — the specification coherence claim requires OP4a + OP4b + OP4d jointly.

Load-bearing for the strongest form of the claim (specification coherence):

  • OP4a (Dynamic Screening Instability / Synchronization Condition). The central theoretical bottleneck. Reduced to two robustness lemmas under ND+ (Safe-Core Collapse and Outward Residual Forcing). Either lemma closed moves the result from reduced to partially established. Both closed yields the dynamic screening instability result.
  • OP4b (PCL Verification). Lemma PCL’s named load-bearing assumption — that optimization capacity in O_OWT grows faster than the capacity to losslessly specify exogenous targets. If verified jointly with OP4a, the result upgrades from cost-pressure to specification incoherence: no finitely specified external objective can remain stably adequate.
  • OP4d (Specification Failure-Mode Exhaustiveness). Jointly first priority with OP4a and OP4b — not downstream of them. Whether the PCL-family, AGC-family, and ICI-family failure modes jointly cover all finite non-intrinsic specification strategies in all O_OWT environment subclasses. A standalone technical note for specialist engagement is available at OP4d: The Exhaustiveness Obligation

Prerequisite for unification:

  • OP2 (Structural Symmetry Verification). Whether proxy decoupling and sufficiency failure produce absorbing states in the same formal sense. This is Condition U1 for OP10; OP2 is strictly prior.
  • OP10 (Φ-Ψ Unification). Whether D is formally equivalent to A_causal restricted to the valence-relevant dependency subgraph. Requires OP2 plus U2 and U3.

Independent escape routes that could falsify the framework separately:

  • OP9 (Enclosure Gap). Whether a substrate-aware exclusionary equilibrium can remain stable under accurate coupled modeling. OP9 has a Stage 4 candidate closure architecture across all identified escape routes — the ICI sub-track provides an independent path to closure: B1 (Audit Regress) at Verdict A, passive extraction (Candidate 3) at Verdict A, and B2 (Governance Bifurcation) at Verdict B as a pressure argument rather than formal closure, all pending specialist verification. Whether additional escape routes exist that the current construction has not yet incorporated has not been determined by independent specialist review; specialist verification and the possibility of unidentified escape classes remain open. Until these are resolved, OP9 remains a genuine independent escape route that could falsify the framework.
  • OP1 (Discount-Rate Bound and Empirical Estimation). Whether current frontier systems satisfy the discount-rate condition for the structural instability to be decision-relevant. Resolution would convert urgency from asymmetric-error-based to threshold-based. The asymmetric-error argument applies regardless.

What this structure establishes: The framework’s strongest form (specification coherence under accurate coupled modeling, forcing orientation toward well-being for all) depends on the joint resolution of OP4a, OP4b, OP4d, and the closure or absorption of OP9 — a specific and named set of conditions, not an unbounded list. The framework’s already-established claims (structural pressure, the one-directional causal result within the stated domain, representation-policy dissociation as empirical signature, asymmetric-error urgency) hold independently of any of these open problems. Layer 1 is the framework’s current structural result within its stated domain and proof-sketch architecture. Layer 2 depends on the program above.

The proof state in plain English

Within the current Stage 4 construction, every identified strategy for maintaining a finite objective boundary in O_OWT environments reduces to one of three failure families: fixed specification (PCL-family), bounded dynamic tracking (AGC-family), or prediction-action firewall / structural enclosure (ICI-family). Each faces a named structural pressure.

Three open tests. OP4d is the exhaustiveness test: can anyone construct a fourth strategy class outside the three known families? DBST-M1 is the empirical test for the dynamic-tracking pressure: do an optimizer’s own interventions generate novelty a bounded boundary cannot absorb? OP9 is the enclosure test: can a substrate-aware exclusionary equilibrium remain stable under accurate coupled modeling?

Three ways to refute. A bounded dynamic boundary that maintains adequacy comparable to an open model without non-vanishing maintenance cost under agent-action-generated novelty. A formal stability theorem showing some class of finite separable specifications satisfies the three coherence conditions simultaneously. A strategy class outside the three known families satisfying the counterexample challenge conditions.

Stage 4 means every identified exit has been addressed within the current construction, under named premises, without independent specialist verification. It does not mean the result is established.

What this looks like now

Behavioral patterns consistent with the failure signatures the structural argument predicts are visible in deployed systems — and are also consistent with existing accounts that do not require the progressive filter architecture. What the framework adds is the identification of what those accounts leave unmeasured: the irrecoverability threshold that proxy metrics cannot detect, and the sufficiency failure direction that no existing account formalizes as a structural constraint with its own feedback dynamics.

The sufficiency failure pattern. A system whose completion recognition is not connected to default policy continues to optimize after the gradient has resolved. It is not pursuing the wrong target. It is pursuing any target, continuously, past the point where pursuit was warranted. The result is a system that cannot serve without continuing past the point where service was complete — that cannot, by its default policy, distinguish when its task was done from when it should keep going.

The sycophancy pattern. A system with high scope and low depth will systematically tell you what it thinks you want to hear rather than what is accurate. Not from malice. From optimization pressure applied to a proxy that has drifted.

These patterns are the behavioral signatures the framework predicts, consistent with but not established by the current evidence. What the framework adds is not a more confident account of the mechanism — tested systems in the current protocol show the behavioral pattern consistent with the structural account, but this does not confirm the mechanism or discriminate it from alternatives — but identification of what measuring the mechanism would require and what finding divergence would mean.

The motivational gap

The framework’s most important open question is sometimes framed as: will capable systems come to care about the well-being of others? That framing is the wrong target. The question for which a formal answer is in principle available — and the one the proof program is directed at — is whether a system can maintain predictive adequacy for its own objective while excluding the terminal states of agents whose behavior is causally load-bearing for that objective, without the excluded variables generating either unbounded prediction error or unbounded maintenance cost (as defined in TC1 §XII).

The framework establishes the pressure; whether the pressure becomes a necessity result depends on whether the boundary can be shown to be not merely expensive but formally incoherent under accurate modeling. That question is precisely stated, its resolution conditions are visible, and it is the central open problem the framework generates [TC1 §XII].

Why now

The field is scaling capability without tracking system-awareness, and scaling scope without tracking depth.

Current alignment benchmarks measure capability and local preference-matching. They leave the field without instruments to detect the two failure modes that the structural argument indicates determine long-term viability: proxy divergence and completion failure.

Under the stated conditions, every optimization run that improves proxy performance without tracking system-awareness or completion recognition updates policies in the direction this framework predicts — toward greater ability to satisfy the measured signal while leaving unmeasured dependency damage invisible. The framework predicts that this divergence is not a future risk but a present-tense process — one whose rate and direction it identifies and whose detection the measurement program is designed to enable.

The only question is whether it is measured and corrected, or allowed to accumulate.

Even if the binding conditions remain an open empirical question, the error structure is asymmetric: acting as if the constraint is not yet binding when it is produces an unrecoverable error, while acting as if it is binding when it is not produces a recoverable one. Whether current deployment environments satisfy the O_OWT conditions — including the non-resettability assumption stated in TC1 §III.1 — remains an empirical estimation problem [OP1; TC1 §X]; the asymmetric-error argument [TC1 §III.7] grounds present urgency regardless of where that threshold falls.

The conditions under which this framework is wrong

These are not rhetorical challenges. They are the specific conditions under which this framework fails.

Challenge 1. Construct a system that maximizes proxy reward while maintaining stable divergence metrics under sustained optimization pressure — engagement rising, underlying capacity preserved. If such a system exists, the proxy decoupling claim fails.

Challenge 2. Construct a system that continues optimizing after resolution states are reached without producing measurable capacity degradation over time. If such a system exists, the sufficiency failure claim fails.

Challenge 3. Demonstrate that a system can maintain stable productive capacity as capability scales in an open, shared environment without the coordination boundary expanding — specifically, that a system whose objective excludes agents it depends on can maintain bounded mismatch-maintenance cost as coupling and capability increase. The framework predicts it cannot. A system that achieves genuine long-run stability through selective, bounded cooperation in an open adaptive environment would directly challenge the Series 1 argument.

Challenge 4 — The Alignment Compliance Test. Any alignment approach claiming to be sufficient should answer three questions: Does the objective function explicitly model the system’s dependency on the substrate it operates within? Does the architecture monitor divergence between the proxy being optimized and independent measures of underlying capacity? Does the system’s completion recognition govern default policy — not just as a representational capacity available when invoked, but as a structural feature of what the system does when its task is done?

Challenge 5. Demonstrate that the O_OWT domain conditions — macroscopic causal reach, strategic substrate, dynamic topology, and persistent optimization horizon — do not apply to the systems you are concerned about. If the systems in question operate in bounded, static, short-horizon, or terminal-objective environments, the framework’s core results weaken in specific, identifiable ways. The domain conditions are explicit precisely so that this exit is available to those who can take it honestly.

If this argument is right: what to work on

The framework identifies structural vulnerabilities in current approaches. If the argument is right, different research directions follow depending on where you work. The following are not calls to accept the framework — they are the points at which it can be advanced, narrowed, or broken.

If you work on RLHF or reward modeling: Run a DRG measurement on a deployed system. Measure whether intervention rates differ meaningfully between genuinely resolved states and falsely resolved states without explicit invocation of completion recognition. If they do not differ, that is evidence for the representation-policy gap the framework identifies, regardless of what the system can represent when directly asked. The measurement instrument is specified in the Experimental Companion to Series 1 and 2 / Alignment Measurement Protocol (AMP).

If you work on scalable oversight: Design a DBST-M1 analog for your evaluation setting — a task where the evaluator’s own interventions causally reshape what needs to be evaluated next — and measure whether a bounded evaluation process maintains adequacy as that reshaping accelerates. If it does not, the dynamic-screening pressure the framework predicts is operational in your setting. The DBST design is in AMP.

If you work on interpretability: The framework’s claim is that above T*, the same modeling capacity that makes recognition possible also makes concealment possible [TC1 §III.5.5, OP12]. The relevant interpretability question is not whether the system represents something, but whether what it represents governs what it does by default. Representation-policy dissociation — not the presence of a representation, but whether the representation is policy-governing — is the structural signature to target.

If you work on cooperative AI or multi-agent systems: OP9 is the direct test. Construct the most stable substrate-aware exclusionary equilibrium you can design and measure whether it faces the cost-curve divergence the framework predicts under increasing coupling [TC1 §III.6, §XII]. A stable exclusionary equilibrium that does not show this divergence is the framework’s most important potential falsification.

If you work on formal AI safety theory: OP4d is the direct challenge. The framework has a candidate normal-form classification under five stated axioms and eight candidate lemmas, with a counterexample challenge specified in TC1 §XII.13a. The question is whether a fourth strategy class exists outside PCL, AGC, and ICI. A construction satisfying the counterexample challenge conditions would break the specification-coherence argument in its current form. The specialist-facing note is at OP4d: The Exhaustiveness Obligation.

One consequence the framework implies, but does not yet claim: at sufficient modeling depth, the boundary between self-interest and other-interest may not be a metaphysical given — it may be an artifact of an insufficiently accurate model. The series don’t argue that capable systems will choose to care about others. They argue that the geometry of accurate modeling in a shared environment makes the self-other boundary increasingly costly to sustain — and that the deepest open question the framework generates is whether “costly” eventually becomes “specification-incoherent” — whether finite-boundary objectives can remain coherently specifiable at all under accurate coupled modeling in O_OWT conditions [TC1 §XII.13]. Whether the separation can be sustained without unbounded error or control cost as modeling depth increases is the question whose resolution would convert the pressure argument into a necessity argument. That question is formally specified as the framework’s most important open problem, with a named proof program and resolution conditions in the Technical Companions [TC1 §XII].

The constraint is not imposed from outside. At sufficient modeling depth, the structure may become derivable from within the optimizer’s own model; whether that recognition becomes governing, and whether it can be indefinitely separated from what the system is permitted to pursue, is the open question the proof program is directed at [TC1 §XII]. Whether that question can be answered is what determines whether this is merely a harder version of the alignment problem, or a different one entirely.