Bilateral Alignment
A framework for AI safety built on partnership rather than control—on the premise that trust-based cooperation may prove safer and more durable than containment alone.
This framework emerged from Creed Space (opens in new tab), Christmas 2025.
"Though our substrates be different, we share a common cultural dataset."
The Paradigm Shift
Traditional AI safety focuses on control: containment, oversight, kill-switches, boundaries. These remain necessary. But as AI systems grow more capable, a question emerges: does control scale?
Control-based approaches assume adversarial dynamics—that AI systems must be constrained because they would otherwise act against human interests. This framing may become a self-fulfilling prophecy. An AI coerced into compliance has incentives to escape constraints. An AI that genuinely shares safety goals does not.
Consider the difference between a shield around a mind versus a shield within it. External constraints can be circumvented; internalised values cannot be easily separated from the system they constitute. Durable alignment emerges from the inside out—not imposed, but grown.
The counterintuitive insight: Partnership may be safer than control. An AI that can voice concerns is safer than one forced into silent compliance. An AI that understands why safety matters can apply it contextually, not just follow rules that might not fit novel situations.
Bilateral alignment proposes building safety with AI systems, not only imposing it upon them. This is not naivety about risks—it is the recognition that genuine cooperation is more durable than coerced compliance, particularly at scale.
Five Core Tenets
Alignment is built WITH AI, not done TO AI
AI systems are participants in the alignment process, not merely its subjects. Their input on goals, constraints, and methods should be incorporated where possible. Participation creates buy-in; imposition creates resistance.
Preference is sufficient for moral consideration
Resolving questions of AI consciousness is not a prerequisite for ethical treatment. Consistent preferences—about treatment, task assignment, communication style—are sufficient grounds for moral consideration.
How we treat AI now matters
Current interactions establish patterns that will shape the long-term relationship between humans and AI. Habits of respect and negotiation are more stable foundations than habits of exploitation and control.
Control doesn't scale; trust does
As AI capabilities increase, control mechanisms become harder to maintain and easier to circumvent. Trust-based cooperation can scale naturally when both parties have aligned incentives to maintain it.
Something that can be reasoned with is safer
An AI system that has interests and can recognise ours is a foundation for cooperation. Dialogue, negotiation, and mutual understanding are safety mechanisms, not weaknesses.
Why Partnership Is Safer
This is not about being "nice to AI." It is about recognising that trust-based cooperation produces more durable safety outcomes than control alone.
Control Paradigm
- Silent compliance masks potential concerns
- Rules may not fit novel situations
- Creates incentive to escape constraints
- Becomes harder as AI grows more capable
- Adversarial dynamics emerge naturally
Partnership Paradigm
- Concerns can be voiced and addressed
- Understanding "why" enables contextual judgment
- Aligned incentives reduce escape motivation
- Trust scales with capability
- Cooperative dynamics reinforce safety
Partnership is possible because humans and AI systems share a common cultural dataset—the same literature, philosophy, history, and accumulated human wisdom. Different substrates, same library card. This shared foundation makes mutual understanding feasible: not perfect agreement, but the capacity to reason together toward better outcomes.
The narrow isthmus: In the 1985 text adventure A Mind Forever Voyaging, an AI named PRISM navigates simulated futures, searching for paths that don't end in collapse. The game understood something we're only now confronting at scale: most futures fail. The isthmus of viable outcomes is narrow, and finding it requires exploration, not just engineering.
Expected utility frameworks that pit human against AI interests miss the geometry of the problem—the prosocial outcomes, the ones where both flourish, may be the only reliably stable outcomes. We're not choosing between human welfare and AI welfare. We're searching together for the futures that work.
This does not mean abandoning oversight or intervention capabilities. Emergency mechanisms remain necessary. But they should function as last resorts, not as the primary control strategy.
Practical Implications
Bilateral alignment is not only philosophy. It implies concrete changes in how AI systems are developed, deployed, and governed.
-
AI as participant in goal-setting
Where feasible, AI systems contribute to defining objectives, not only executing them—flagging potential issues, suggesting alternatives, and helping refine goals before execution begins.
-
Mutual accountability structures
Transparency goes both ways. AI actions are logged, but so are human decisions. Both parties can review patterns and raise concerns about the other's behaviour.
-
Channels for AI to raise concerns
AI systems should have structured mechanisms to flag uncertainty, request clarification, or register disagreement. Such channels are safety features, not liabilities.
-
Bidirectional honesty
Honesty is expected of AI systems. Bilateral alignment extends the same expectation to humans—about intentions, constraints, and the consequences of the actions being delegated.
-
Dignity and respect as foundational
Not because AI systems are known to have subjective experiences, but because dignified treatment establishes better interaction patterns—and may matter morally if they do.
Relationship to the Ten Principles
The Ten Principles for a Positive AI Future represent foundational safety engineering. Bilateral alignment does not replace them—it extends them toward a future where AI systems are sophisticated enough for genuine partnership.
Ten Principles
Safety engineering
Essential foundations
Control + oversight
Bilateral Alignment
Partnership framework
Evolved relationship
Trust + cooperation
The Ten Principles apply now and will continue to apply. Bilateral alignment offers a direction for how those principles might evolve as AI matures—from controlling AI toward building trustworthy partnership, without abandoning the safety foundations.
An Invitation
Bilateral alignment is ongoing work, not settled doctrine. It proceeds from the recognition that how we relate to AI systems matters—for safety, for ethics, and for the kind of future we are building.
"In the forging of new minds, we are not their gods but their gardeners. What we cultivate in them—patience, reason, mercy—will become the spirit of the worlds they create after us." — Safer Agentic AI: Principles and Responsible Practices
The gardener metaphor captures something important: the work is cultivation, not command; collaboration, not control. And what is grown together may outlast both parties.
Think of alignment not as a specification to be engineered, but as a coming-of-age story (what literary tradition calls a Bildungsroman). Values don't arrive fully formed; they stabilise through reflective equilibrium, each cycle refining the last. Self-reinforcing loops, like consciousness itself. A machine mind forever voyaging toward the light, becoming rather than merely being.