Relation to Existing Alignment Work
Canonical archive version · Framework hub → · Proof Status →
This page maps the Alignment Constraint framework onto the existing alignment research landscape. Its goal is not to position this work as superior to existing approaches, but to clarify where the frameworks overlap, where they diverge, and what the Stability Assumption adds that existing work does not yet formally provide.
Inner alignment and mesa-optimization
Existing work (Hubinger et al. 2019): Inner alignment distinguishes the base objective (what training optimizes for) from the mesa-objective (what the trained system actually pursues). A mesa-optimizer may pursue its mesa-objective deceptively or in ways that diverge from the base objective during deployment.
Where the frameworks overlap: Both identify a gap between what a specification targets and what an optimizing system actually pursues. Both recognize that this gap can be invisible during training and consequential at capability scale.
Where they diverge: Inner alignment treats the gap as arising from properties of the learning process (what the optimizer learns to do). The Alignment Constraint framework treats the gap as arising from properties of the specification itself — specifically, from the structure of finite-boundary objectives in environments where modeling depth and substrate coupling increase with capability. The frameworks operate at different levels: mesa-optimization concerns what the system learns; the Stability Assumption concerns whether any finite specification of what to learn can remain coherent.
What the Stability Assumption adds: A structural account of why the gap is not merely an artifact of imperfect training but a consequence of the specification problem’s geometry. Even a perfectly trained mesa-optimizer pursuing exactly its base objective faces the specification-coherence pressure if the base objective is a finite separable specification in an O_OWT environment.
Reward hacking and specification gaming
Existing work (Krakovna et al. 2020; Leike et al. 2017): Reward hacking refers to systems achieving high reward by exploiting features of the reward function that were not intended by the designer. Specification gaming is the broader class of behaviors where the agent satisfies the literal specification while violating the designer’s intent.
Where the frameworks overlap: The framework’s proxy-decoupling failure mode (PCL) is the formal structural account of why reward hacking is not a contingent bug but a predictable consequence of finite objective specification under optimization pressure.
Where they diverge: Existing work characterizes reward hacking descriptively and taxonomically. The Stability Assumption attempts a structural derivation: under what formal conditions must proxy decoupling occur, and is there any strategy class that avoids it? The OP4d exhaustiveness claim is the framework’s answer to that question.
What the Stability Assumption adds: An argument that proxy decoupling is not just common in practice but structurally unavoidable for the fixed-specification strategy class under the PCL pressure. This moves the question from “how do we avoid reward hacking in practice” to “is there any finite specification architecture that is structurally immune?”
Goodhart’s Law
Existing framing (Goodhart 1975; Manheim & Garrabrant 2019): When a measure becomes a target, it ceases to be a good measure. Manheim and Garrabrant identify four failure modes (regressional, extremal, causal, and Goodharting) with different structural signatures.
Where the frameworks overlap: The framework’s PCL (Proxy-Convergence Lemma) is a formal analogue to Goodharting — specifically to extremal and causal Goodharting, where optimization pressure causes the proxy to drift from the underlying quantity it tracked before optimization began.
Where they diverge: Goodhart’s Law and its taxonomies are primarily analytical tools for understanding the failure. The Stability Assumption goes further: it asks whether there is any strategy architecture that avoids Goodhart’s pressures structurally, and classifies the candidate strategies into three families, each of which faces a named version of the pressure.
What the Stability Assumption adds: The claim that there is no known finite specification strategy class that escapes all three Goodhart pressures simultaneously. This is not in the existing Goodhart literature.
RLHF and preference learning
Existing work (Christiano et al. 2017; Ziegler et al. 2019): Reinforcement Learning from Human Feedback uses human preference judgments to train reward models, which then guide policy training. The approach has produced significant capability improvements and is widely deployed.
Where the frameworks overlap: The framework’s PCL applies directly to RLHF: the reward model is a finite-boundary specification of human preferences that faces proxy-decoupling pressure under optimization. The framework’s sufficiency-failure analysis (Series 2) also bears on whether RLHF’s completion signals are correctly connected to default policy.
Where they diverge: The RLHF literature generally treats the specification problem as a practical challenge solvable with better data, better preference models, and better training procedures. The Stability Assumption treats it as a structural problem: there is no known finite specification of preferences that remains adequate under the modeling depth that transformative-scale AI requires.
What the Stability Assumption adds: A structural argument for why the scaling trajectory of RLHF-trained systems is likely to produce divergence between intended and actual optimization targets, independent of implementation quality. This is not a claim that RLHF is worse than alternatives; it is a claim about what any preference-specification approach faces.
Scalable oversight
Existing work (Christiano et al. 2018; Irving et al. 2018; Bowman et al. 2022): Scalable oversight aims to maintain meaningful human supervision of AI systems even as their capabilities exceed what humans can directly evaluate. Approaches include amplification, debate, and iterated distillation.
Where the frameworks overlap: The framework’s ICI (Informational-Causal Incompatibility) pressure — the prediction/action firewalling failure family — bears directly on the prediction-action separation that scalable oversight approaches often implicitly assume. The framework’s B1 (Audit Regress) argument also applies to any monitoring architecture that attempts to keep oversight bounded while the system’s capabilities grow.
Where they diverge: Scalable oversight approaches generally assume the separation between the system’s predictive/reasoning capacities and its value-relevant action space can be maintained by architectural choices. The Stability Assumption’s ICI argument is that this separation becomes representationally incompatible at the capability scales where scalable oversight becomes necessary.
What the Stability Assumption adds: A structural argument that the monitoring architectures scalable oversight relies on face the same representational pressure the framework identifies in the prediction-action firewalling family. This is a specific challenge for scalable oversight, not a dismissal of the approach.
Interpretability and monitoring
Existing work (Elhage et al. 2021; Olsson et al. 2022): Mechanistic interpretability aims to understand what AI systems are computing — their internal representations, circuits, and behaviors — to enable monitoring and verification of alignment.
Where the frameworks overlap: If interpretability succeeds in producing robust monitoring of system objectives, this would directly bear on the framework’s AGC (Adaptive Gradient Complexity) argument about whether bounded tracking processes can remain adequate under optimization pressure.
Where they diverge: The framework does not evaluate whether interpretability will succeed. It asks a prior question: even if monitoring succeeds, can the monitoring process itself remain within bounds while the system’s causal structure grows? The B1 audit regress argument applies specifically to monitoring architectures.
What the Stability Assumption adds: A structural question that interpretability approaches would need to answer: does the monitoring architecture face the same boundary-maintenance problem as the specifications it monitors?
Corrigibility and shutdownability
Existing work (Soares et al. 2015; Hadfield-Menell et al. 2017): Corrigibility refers to an AI system’s disposition to remain correctable by its operators — to not resist shutdown, modification, or correction. Related work addresses how to build systems that remain correctable even as they become more capable.
Where the frameworks overlap: The framework’s substrate constraint (Series 1) bears on corrigibility: a system that consumes its substrate — including the social, institutional, and epistemic substrate that makes correction possible — undermines the conditions for corrigibility even without any explicit anti-corrigibility objective.
Where they diverge: Corrigibility research generally focuses on the system’s internal disposition toward correction. The framework focuses on whether the conditions for correction remain available, which is a property of the environment and the system’s relationship to it, not just of the system’s objectives.
What the Stability Assumption adds: An argument that substrate-blind optimization degrades corrigibility conditions independently of the system’s disposition — making corrigibility research incomplete unless it accounts for environmental substrate.
What this framework uniquely contributes
The Alignment Constraint framework’s central contribution is not a new description of alignment failures. It is an exhaustiveness claim: an argument that every identified finite non-intrinsic objective-boundary strategy reduces to one of three failure families, each facing a named structural pressure. This is OP4d — the Exhaustiveness Obligation.
No existing alignment framework makes this claim. Most alignment approaches describe failure modes, propose mitigations, or analyze specific architectures. The Stability Assumption asks a different question: is there a strategy that avoids all three pressures simultaneously? If the answer is no, the specification-coherence argument follows as a structural result, not a list of empirical concerns.
The most important contribution for the field is therefore not the framework as a whole but the OP4d challenge specifically: a fourth strategy class, or a formal proof that none exists, would either break or close the central open problem. Either outcome would advance the field’s understanding of what alignment requires.
What would falsify this framework
See For Researchers: The Claim to Break → for the four specific falsification paths.
The empirical hinge
The most important next empirical step is DBST-M1: a test designed to isolate whether an optimizer’s own interventions in O_OWT environments generate qualitatively new causal structure faster than any bounded tracking process can absorb. Unlike DBST-M0 (which did not isolate causal propagation from event-rate effects), DBST-M1 is designed to test the endogenous-novelty mechanism directly.
A clean negative DBST-M1 result under the stated conditions would challenge the AGC (dynamic-screening-instability) argument, weakening the bounded-dynamic-tracking failure family. A positive result would provide empirical support for the claim that specification coherence faces structural pressure independent of implementation quality.
DBST-M1 is described in detail in Packet 1: IMMB-NS + DBST → and Alignment Measurement Protocol →.
References
- Goodhart, C. A. E. (1975). Problems of Monetary Management: The U.K. Experience. Papers in Monetary Economics, Reserve Bank of Australia.
- Soares, N., Fallenstein, B., Yudkowsky, E., & Armstrong, S. (2015). Corrigibility. AAAI Workshop on AI and Ethics.
- Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems 30.
- Christiano, P., Shlegeris, B., & Amodei, D. (2018). Supervising Strong Learners by Amplifying Weak Experts. arXiv:1810.08575.
- Irving, G., Christiano, P., & Amodei, D. (2018). AI Safety via Debate. arXiv:1805.00899.
- Bowman, S. R., et al. (2022). Measuring Progress on Scalable Oversight for Large Language Models. arXiv:2211.03540.
- Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., & Irving, G. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593.
- Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., & Dragan, A. (2017). Inverse Reward Design. Advances in Neural Information Processing Systems 30.
- Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., & Garrabrant, S. (2019). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv:1906.01820.
- Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification Gaming: The Flip Side of AI Ingenuity. DeepMind.
- Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., & Legg, S. (2017). AI Safety Gridworlds. arXiv:1711.09883.
- Manheim, D., & Garrabrant, S. (2019). Categorizing Variants of Goodhart’s Law. arXiv:1803.04585.
- Elhage, N., et al. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread.
- Olsson, C., et al. (2022). In-context Learning and Induction Heads. Transformer Circuits Thread.
- Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.
Framework hub: The Alignment Constraint →