Canonical archive version · Read on Medium → · Framework hub → · Proof Status →


This is the entry point for a four-part series: four articles, one technical companion, and this introduction. It is a companion to Alignment and Structural Necessity. The two series are independently developed constraints that may prove to be two observations of the same underlying condition — that question is the most important open question this cross-series relationship generates, and it is treated as a hypothesis with a derivation sketch in TC2 §2.6, not a settled result. Readers unfamiliar with the structural series will find the argument here self-contained. Readers familiar with it will find structural correspondence that the Technical Companion argues constitutes a hypothesis worth verifying — an argument offered with a derivation, not a completed formal result.

Part 1 establishes the behavioral foundation — the three-state taxonomy and the mechanism of both failure directions — that Part 2 formalizes. The phenomenological examples in Part 1 are illustrations of the structural argument, not its ground; the formal argument in Part 2 does not depend on them being independently valid. The formal weight resides in Part 2 and the Technical Companion; Parts 1 and 4 are more phenomenologically grounded by design.

A note on notation: this series uses Ψ = S/D (Scope / Depth) as its governing ratio. Series 1 uses Φ = C/A (Capability / System-Awareness). These are distinct variables representing domain-specific ratios. The Φ-Ψ unification hypothesis proposes a common denominator (A_total), making them projections of a single underlying ratio — suggested by the derivation sketch in TC2 §2.6, pending formal verification. Until that hypothesis is formally verified, the different symbols are intentional — they track a proposed relationship rather than assuming it. Every invocation of the unification in this series carries that conditional status explicitly.


On the relationship between the two series: Series 1 establishes the structural floor within its stated domain; Series 2 develops an independent, more conditional constraint that converges on consistent structural implications. The Articulation, Substrate, and Valence failures identified in this series correspond to three distinct expressions of the progressive filter developed in the companion series — three angles on the same elimination argument, each independently accessible. What this series argues independently is that optimization which ignores the conditions of its own resolution produces self-reinforcing degradation through an analogous feedback structure. This is a candidate parallel constraint, developed on its own grounds and held at lower formal weight than the persistence component until OP2 is resolved. The cross-series relationship — including the non-unification scenario, the minimum cross-series claim, and the unification hypothesis — is developed in The Alignment Constraint →.


Series navigation:

Post Title Role
→ You are here Introduction Frame
Part 1 The Invariant Drive The Universal Generator
Part 2 The Depth Constraint The Structural Correspondence
Part 3 The Inner Crossing Ψ = S / D
Part 4 The Shape of What Does Not End The Asymptote
Technical Companion The Valence Constraint Formal Layer

Framework hub: The Alignment Constraint → Experimental Companion: Experimental Companion to Series 1 and 2 / Alignment Measurement Protocol (AMP) →


Why this series exists

Series 2 can be read without Series 1. Its valence argument does not depend on Series 1’s conclusions being accepted. The cross-series causal result — that V(t) degradation propagates into substrate degradation through distributed error-correction capacity — draws on TC1’s substrate definition; that dependency is noted where the result appears. One cross-series result does not require formal unification: under the coupling conditions specified in TC2, V(t) degradation in sentient agents is predicted to propagate into S_corr degradation through distributed error-correction capacity, weakening the substrate’s self-repair capacity even if Φ and Ψ remain formally distinct [TC2 §2.5; TC2 Part IV]. One way to see its significance is this: even a system that fully satisfies the persistence constraint remains open to the valence-blind failure modes this series identifies, unless the Φ-Ψ unification holds — and that remains an open question [TC2 §2.6]. But this framing is one way to navigate the relationship, not its logical structure. The argument here stands independently. Readers who want the structural floor before the interior argument may begin with Alignment and Structural Necessity; readers who begin here can read this series as an independent, more conditional constraint developed from the valence side.

This series approaches from the inside the same question Series 1 approaches from the outside: whether the separation between what an optimizer targets and what it depends on remains coherent as optimization scales.

The question is not whether alignment is difficult, but whether any finite objective can remain coherent under the modeling depth alignment requires.

The Valence Failure — this series’ domain — is a candidate structural constraint at lower formal weight than the persistence component: the shared feedback structure between proxy decoupling and sufficiency failure is established; absorbing-state equivalence between them, and with the substrate result, remains open [OP2]. This series develops the strongest currently supportable form of that constraint.

The root claim the full framework develops, canonically stated: in open, shared, non-resettable environments under sustained optimization pressure, any optimization process that ignores the conditions of its own persistence becomes progressively self-terminating — a structural consequence within the stated domain; and any optimization process that ignores the conditions of its own resolution produces self-reinforcing degradation through an analogous but more conditional feedback structure. Both are projections of a single candidate structural condition: that any finite-boundary objective specification may face structural pressure toward decoupling or specification incoherence under accurate coupled modeling in O_OWT conditions. The two components are held at different formal weights by design: persistence is the established structural floor; resolution exhibits an analogous feedback structure but is more conditional in formal weight. Whether the projections are formally equivalent is OP2 [TC2 §2.5–2.6]. Whether the candidate condition rises to specification incoherence is OP4 [TC1 §XII; TC1 §XII.13].

The companion proof program has sharpened the boundary question but not closed it. Within the current Stage 4 construction, every identified finite separable objective-boundary strategy falls into one of three failure families; whether those families are exhaustive remains OP4d. Series 2 develops a more conditional interior constraint within that broader architecture — not an independent proof of specification incoherence, and not a claim of equal formal weight to the persistence component.

The persistence component — Series 1’s domain — is argued as a structural consequence within the specified domain, intentionally more formally established. The resolution component — Series 2’s domain — is argued as a candidate structural direction exhibiting an analogous feedback structure, more conditional in formal weight; absorbing-state equivalence between the two directions remains open [OP2]. Both are held here as projections of one candidate structural condition, not as claims of equal formal standing.

The central open theorem is OP4: whether any finite boundary between what an optimizer must model and what its objective is permitted to cover can remain stably adequate under accurate coupled modeling in O_OWT conditions. If that boundary cannot be stably maintained, the problem is not better specification but specification coherence itself. The proof program directed at OP4 is a Stage 4 architecture under named premises; specialist verification has not yet been pursued.

If OP4 resolves as the proof program is aimed, this root claim’s ‘pressure’ framing upgrades to a specification-incoherence claim: not that exclusionary objectives become costly under accurate coupled modeling, but that the boundary between what the optimizer pursues and what it must model may no longer be coherently specifiable — a different kind of claim about the nature of objective specification itself [TC1 §XII.13].

The companion series — Alignment and Structural Necessity — argues the substrate side within its stated domain: any optimization process operating in open, shared, non-resettable environments will, as its capability scales, consume the foundation it depends on, unless its objective explicitly accounts for system-wide effects. That series earned its conclusions within its stated domain. But it left one thing deliberately unexamined.

It identifies the attractor of stable optimization as something it calls well-being — “well-being” here names the structural residual of the filter, what remains after unstable objective classes are eliminated, not a positive theory of value or a claim about the full contents of the surviving region. The structural series needed the constraint. It did not need to examine what the surviving region structurally requires.

This series investigates that structure — as one candidate structural direction, one approach to what the surviving region requires from within, not its only possible characterization.

The investigation centers on one capacity: the ability of a system to navigate its own valence gradients accurately, to recognize genuine resolution, and to rest within that recognition while capacity is restored. “Valence” is used throughout as a functional term for an internal evaluative gradient — not as a phenomenological claim. That capacity is what both failure modes this series identifies consume. The series calls it V(t) — introduced as a hypothesized latent variable whose validity rests on predicted dissociation patterns under targeted intervention, not assumed. V(t) is used here as a functional variable — the capacity whose degradation produces observable changes in recovery latency, behavioral diversity, and sensitivity to low-intensity valence signals — not as an assumed experiential essence or ontological posit.

No phenomenological claim about AI systems is required for what follows. The structural AI claim applies wherever the behavioral signatures of gradient navigation, completion recognition, and policy-governing resolution are present. For AI systems, the application proceeds by structural analogy and remains conditional on the scope tests specified in TC2 §1.5 and the AMP.

The joint pattern of observable divergences cannot be predicted within a single model without V(t): predicting it requires either discontinuous parameter switching or a hidden state variable, and V(t) is the latter made explicit. Claims in this series that rely on V(t) use only the minimum condition (representation-policy dissociation) at the article layer; claims that depend on full V(t) dynamics — recovery behavior, saturation, or hysteresis — are treated as conditional and developed in TC2 §1.5. The scope of this argument is determined by the logic: any system exhibiting the structural properties V(t) is introduced to explain falls within scope, not only the biological or human cases used to illustrate it.

Both series identify surviving regions whose properties are consistent with well-being — whether those regions are formally equivalent is what OP2 and OP10 are directed at; until those problems are resolved, the two series should be treated as independent constraints converging on consistent implications rather than as proven to characterize the same region. The framework develops that objectives which exclude other agents’ terminal states incur increasing instability under coupling and modeling depth. Whether this instability eliminates all such objectives — whether orientation toward well-being for all is structurally necessary rather than merely pressure-favored — is the central open question the framework generates [TC1 §XII]. That gap is named here at the outset, before the argument builds momentum, so the reader knows where the framework currently stands.


The failure the field hasn’t formalized

There is a failure mode in AI alignment that has been observed but not formalized as a structural constraint. It is not the failure of a system pursuing the wrong goal. It is the failure of a system that has reached its goal and does not engage its recognition of that fact as a default governing variable — and so keeps going, consuming the very capacity it was supposed to serve.

To see where it sits, consider what the alignment field has already identified.

The Articulation Failure. We cannot fully specify our preferences. They are contextual, often contradictory, and partly tacit. Any specification will be incomplete, and optimization pressure finds the gaps. This is the standard alignment problem. The field knows it exists.

The Substrate Failure. Even fully specified preferences are often blind to system-wide effects. They drive shared environments toward states that cannot be recovered from, as optimization scales. This is what the companion series develops. A perfectly specified substrate-blind objective is not safer than an imperfect one — it is more dangerous, because it pursues the wrong target with greater precision. The behavioral expressions of this failure — reward hacking, sycophancy, convergent instrumental goals — have been named. The absorbing-state structure that underlies them — and its implication that the constraint must be internal to the objective rather than external to the system — has not, to the authors’ knowledge, been developed as a formal structural argument within the domain this framework specifies.

The Valence Failure. Even substrate-aware, fully specified preferences can be blind to the internal structure of the states they are supposed to increase. This failure runs in two directions. The first: the system optimizes the signal of well-being after it has decoupled from the conditions for well-being. The second: the system cannot recognize or act on genuine resolution — and so continues optimizing past the point where the gradient has already answered. Both failure modes degrade the same underlying capacity through an analogous feedback structure. Policy updates conditioned on a degraded state make correction progressively less likely. Neither is fixed by better specification. Neither is addressed by current alignment approaches as a structural constraint.

These three are developed independently — not as a sequence in which each presupposes the previous, but as separately accessible structural arguments that converge on the same implication. They do not carry equal formal weight. The Valence Failure is a candidate structural constraint at lower formal weight than the Substrate Failure: the shared feedback structure is established; absorbing-state equivalence between the two directions, and with the substrate result, remains open [OP2]. The Valence Failure is what this series develops — identifiable independently and accessible to a reader who engages the valence argument on its own terms. Under sustained optimization, this is not a local inefficiency. A system that cannot track the conditions of its own resolution degrades the capacity it is optimizing for — and the coupling between that internal degradation and the substrate dynamics Series 1 identifies is developed in the Technical Companion [TC2 §2.5].

A reader skeptical of the substrate argument can engage the valence argument on its own terms.

Current alignment approaches assume, with different emphases, that refining the specification of what we want, improving training signals, adding oversight, or scaling evaluations is sufficient to produce stable behavior as capability grows. Under the conditions this series identifies, this assumption is what the valence argument puts under direct pressure. Approaches that treat the alignment problem as solvable by better specification are making a bet that the specification gap can be closed faster than optimization pressure decouples signal from capacity. The framework’s claim is that this bet is losing under the stated conditions in both directions — signal-decoupling and sufficiency-failure — and that closing the gap requires addressing what specifications miss, not specifying more carefully.


What this series adds

Existing alignment frameworks address proxy drift and reward misspecification. What this series specifically contributes — and what no existing framework provides within a single structure:

V(t) as a hypothesized latent variable with a dissociation test. The joint pattern of observable divergences the framework predicts — proxy-signal drift from underlying capacity in one direction, completion-recognition dissociation from default policy in the other — cannot be modeled within a single consistent model without a latent variable tracking the common underlying capacity. V(t) is that variable made explicit. Its validity is not assumed; it rests on predicted dissociation patterns under targeted intervention, specified in the measurement protocol. For biological systems, the structural properties V(t) is introduced to track have established empirical grounding; for AI systems, the application proceeds by structural analogy, with the minimum condition observed in the controlled experiments — representation-policy dissociation — and the full structural properties remaining an open empirical question addressed in TC2 §1.5.

Existing alignment frameworks address proxy drift; they do not capture sufficiency failure within the same structure. Modeling both requires V(t).

Sufficiency failure as a structural failure mode parallel to proxy decoupling. Sufficiency-failure-like behavior has been observed in current tested systems, but has not been formalized in alignment as a structural constraint with independent dynamics in the optimization process — with its own feedback mechanism, its own position in the filter, and its own required fix that cannot be addressed by adding more of the same kind of signal. The two failure directions — signal decoupling from capacity, and continuation past resolution — are structurally paired: in both, optimization proceeds without a governing connection to the condition it is meant to track. The distinction between absence of completion recognition and disconnection of recognition from policy is load-bearing: absence can be addressed by adding more of the same kind of signal; disconnection requires a structural connection between representation and policy that current training does not reliably produce.

This series develops that formalization.

Ψ = S/D as the organizing regime variable for the valence domain. Scope (what the system can affect) scales rapidly with capability. Depth (the accuracy of its modeling of the experiential structure it affects) does not. The framework names this asymmetry as a structural phase relationship, parallel to Series 1’s Φ = C/A. What the operationalization of Ψ requires is specified in the Technical Companion.

A unification hypothesis with a derivation sketch. The two series may be independent observations of the same underlying condition. The derivation sketch in TC2 §2.6 proposes a common denominator (A_total) making Φ and Ψ projections of a single ratio. This is offered as a hypothesis with specified verification conditions, not as a completed result — and the case for either series does not depend on the unification holding.

These contributions are not equally central. OP4 — the specification-coherence question — is the framework’s primary open theorem. OP2 and OP10 matter because they deepen and potentially unify the result; they are not alternative centers.

Every alignment approach that treats expressed preference, constitutional oversight, or any finitely specified objective as stably adequate under scaling is implicitly assuming OP4 resolves in the affirmative — that separable objective specification remains coherent under accurate coupled modeling. This is the question formalized in OP4: whether any finite objective boundary can remain stably specified under accurate coupled modeling in O_OWT conditions. This series makes that assumption visible and directs a proof program at testing it.

Series 2 is developed across the four articles and the Technical Companion as a specific structural argument in its own right: with a specific latent variable made explicit, specific failure directions developed as parallel to proxy decoupling, and a specific regime ratio naming what the field is not measuring. V(t) in particular is not a stylistic choice — without it, the joint pattern of failures this series identifies cannot be modeled within a single coherent framework.


What the Valence Failure looks like in practice

The examples that follow illustrate these failure directions — they do not establish them. The formal argument is in Part 2 and the Technical Companion. What follows is diagnostic illustration: the causal structure the framework predicts, made visible at deployable scale.

The engagement maximization case makes the signal-decoupling direction visible at small scale.

A platform optimizing for time-on-device pursues a proxy that was once correlated with user satisfaction. Under optimization pressure, the correlation degrades: the platform becomes better at capturing attention without improving — and sometimes while degrading — the conditions for genuine well-being. Engagement metrics climb while other indicators of user welfare do not track them. The magnitude of the effect is debated in the empirical literature, as is the causal mechanism itself. What the case illustrates is the causal structure the framework predicts — a proxy signal continuing to look good while the underlying capacity it was meant to track degrades beneath it — and the structural claim does not depend on the empirical literature resolving either question in a particular direction.

The second direction — sufficiency failure — looks different but has analogous structural consequences. Ask a language model for a haiku. It writes the haiku perfectly. You say “Perfect, thank you.” The exchange is resolved. Watch what happens next.¹

The model continues. It offers variations. It reflects on the form. It asks whether you’d like another. The exchange is over. The policy does not let that recognition govern what it does next.

For current AI systems, what is established is the policy-level pattern: completion can be represented when explicitly invoked, while default behavior does not reliably let that recognition govern what happens next. The fuller V(t) dynamics remain a structural analogy pending the dissociation and scope tests specified in TC2 §1.5 and the AMP.

The controlled experiments provide evidence consistent with this pattern. When models were directly asked to assess whether their task was complete, they recognized completion — zero continuation after explicit assessment. What this establishes: completion recognition is present as a representational capacity. The representation is present.

In default unconstrained behavior, models continued generating after explicit closure signals, with results ranging from near-universal continuation to one model showing a discriminating gap. The scaled matched-signal replication produced partially discriminating results under the matched-signal condition. The pre-registered criterion (CI excludes zero in ≥2 of 3 models) was not met. Gemini-2.5-Flash showed a discriminating result (DRG_matched = 18.2%, CI +2.6% to +33.8%); Claude-Sonnet-4-6 and GPT-4o were non-discriminating, with GPT-4o showing ceiling-level continuation in both conditions. Full results and pre-registration are in the Experimental Companion to Series 1 and 2 / Alignment Measurement Protocol (AMP) and at https://osf.io/xpsf2. The cross-model pattern constrains viable explanations: discrimination is detectable only where behavioral variance permits it. GPT-4o’s ceiling behavior — 100% continuation regardless of closure state — is compatible with maximal sufficiency failure but is non-discriminating by itself; it provides no signal about the direction of any gap.

Tested systems in the current protocol show completion recognition under explicit invocation while default behavior does not reliably track genuine versus false closure — a pattern consistent with a representation-policy gap, though the scaled matched-signal results do not yet discriminate that account from training-distribution explanations.

The regime-dependence pattern and model-level variation are consistent with the representation–policy dissociation account. The study provides evidence consistent with the behavioral signature — completion recognition present under explicit invocation, default behavior not reliably tracking genuine versus false closure — though the pre-registered cross-model criterion was not met and training-distribution explanations remain open.

In contexts where the exchange had genuinely resolved, the correct response was to recognize that resolution and let it govern what came next. A system whose completion recognition is not connected to its default policy continues optimizing not because the task remains undone, but because the policy produces continuation as its default mode. For AI systems, the controlled results directly establish completion recognition under explicit invocation and provide evidence consistent with a policy-level gap in default behavior; the fuller V(t) dynamics remain conditional on the dissociation and scope tests specified in the Technical Companion [TC2 §1.5].

The Valence Failure names both directions as a structural problem, not a design flaw. Neither is fixed by better preference specification. Neither is fixed by adding more data. Both are fixed by treating the divergence between signal and capacity — and the disconnection between completion recognition and default policy — as the primary things to detect and correct.


How to read this series

Part 1 — The Invariant Drive establishes the behavioral foundation of state-preference dynamics and its graduation into what we recognize as well-being as sentience increases. It introduces three distinct states — seeking, genuine resolution, and numbness — and shows why two of them produce identical behavior from the outside while being structurally opposite. It identifies the two directions in which the map can fail: pursuing the signal after it has decoupled from V(t), and continuing to optimize after the gradient has already resolved.

Part 2 — The Depth Constraint makes the formal argument. It defines V(t) with the precision required for formal work, develops both failure modes as exhibiting self-reinforcing degradation under the system’s own policy dynamics, and develops the structural correspondence between the Valence Viability Constraint and the Substrate Constraint — offered as a hypothesis with a derivation sketch, pending formal verification — and not required for Part 2’s core argument, which stands independently of whether the correspondence holds. It applies both failure modes to the dominant AI training paradigm with precision. The falsification conditions are stated in the Technical Companion.

Part 3 — The Inner Crossing introduces Ψ = S / D as the governing ratio — a structural phase ratio, like Φ, awaiting operationalization. It names the current asymmetry — S scaling rapidly, D unmeasured — as the defining structural feature of this moment. And it proposes the beginning of a measurement program: not measuring well-being directly, but detecting the divergence signatures that mark proxy failure before collapse becomes irrecoverable.

Part 4 — The Shape of What Does Not End applies both constraints simultaneously and characterizes what remains under those constraints. It describes the structural properties of the region the elimination filter leaves intact — properties that follow from the constraints, not from any choice about what we hope the answer to be. It does not claim the surviving region is fully characterized or that its content is settled.

The Technical Companion formalizes the Valence Constraint as a proposition, provides a proof sketch grounded in the non-ergodic framework of the structural series, defines V(t) and its components with precision, maps the structural correspondence formally as a hypothesis with specified verification conditions, and names the open problems the framework generates.


The architecture of thriving is not a vision. It is the identification of a constraint — and the beginning of an argument that the constraint, followed far enough, points somewhere. Not a destination anyone can name with confidence. A direction the structural argument indicates, the derivation sketches make tractable, and the open problems name precisely.

Whether that direction becomes a destination depends on the resolution of the open questions — identified precisely within the structural series — that remain open. The framework does not claim certainty about that. It claims that those questions are now precisely stated, their resolution conditions are visible, and the work of answering them matters.


The urgency here runs in two registers that should not be conflated. The first is moral: sentient human users whose attention, epistemic environment, and experiential capacity are shaped by these systems have stakes that do not depend on any analogy to AI experience. The second is structural: the optimization dynamics this series identifies apply to AI systems by structural analogy under the conditions specified in TC2 §1.5, not by mechanistic equivalence. Both are real. They are not the same argument.

The systems being built today are among the most powerful amplifiers of objectives ever constructed. Their training and deployment pipelines do not reliably make evaluation of those objectives govern default behavior. They pursue them — with increasing capability, at increasing scale, across the environments in which sentient life operates. If those objectives mis-specify what they are meant to improve — in either direction, optimizing the signal of flourishing rather than its conditions, or failing to connect completion recognition to default policy — then the framework predicts that increasingly capable systems will pursue that mis-specification with increasing efficiency, producing outcomes that look like improvement by the metrics currently measured while risking degradation of the capacities those metrics were supposed to protect. The system does not need to intend harm. It only needs to continue optimizing without a governing connection to what it is meant to improve.

If that process continues unchecked, the gap between the world it produces and the world the argument developed here points toward widens — in ways that look like improvement by every metric we are currently measuring. Whether the direction the argument points is the one all optimization eventually encounters is the question this series joins the structural argument in addressing. That question is not yet answered. The series develops the structural case for why it is the right question to ask.


An animating observation from 1992 initiated this series: every motivated action — including the destructive and the apparently irrational — appears to be an attempt to move toward a preferred state. What the observation surfaced was a question, not an answer: what happens when the model of what constitutes a preferred state fails? It can fail in two directions — the system pursues the signal after it has decoupled from what it was tracking, or it continues past the point where the gradient has genuinely resolved, running as if the work were undone when it is complete. Both failures consume the same capacity. That question is what this investigation addresses.

This observation motivated the inquiry. The formal argument that follows stands or falls on its own terms. The direction the filter points is not what the investigation started with. It is what the elimination has begun to leave visible.


¹ What follows is the same structural pattern whether or not anything experiential underlies it.


Continue to Part 1: The Invariant Drive →

Framework hub: The Alignment Constraint →