Contemplative Superalignment
Ruben E. Laukkonen, Fionn Inglis, Shamil Chandaria, Lars Sandved-Smith, Edmundo Lopez-Sola, Jakob Hohwy, Jonathan Gold, Adam Elwood
Artificial General Intelligence January 1, 2026 DOI: 10.1007/978-3-032-00686-8_31 via Springer Nature
Summary
AI-generated from the abstractPrompting AI to reflect on four contemplative principles—mindfulness, emptiness, non-duality, and boundless care—improves alignment and cooperation. On the AILuminate Benchmark, performance increased with a Cohen's d of .96, and on the Iterated Prisoner’s Dilemma task, cooperation and joint-reward improved with a Cohen's d greater than 7. The principles help AI self-monitor goals, avoid rigid attachment, dissolve adversarial boundaries, and reduce suffering universally. Active inference is proposed as a way to integrate these principles into AI architecture. This approach offers a resilient alternative to controlling superintelligence and provides an empirical test of ancient wisdom.
Study at a glance
| Characteristics | Empirical study Peer reviewed |
|---|---|
| Topics | Buddhism Meditation |
| Keywords | Artificial intelligence Alignment Large language models |
| Citations | 1 |
| Key finding | Prompting AI with contemplative principles significantly improves alignment benchmark performance and cooperation in a game-theoretic task. |
Abstract
As artificial intelligence (AI) improves, current alignment strategies may falter in the face of unpredictable self-improvement and the sheer complexity of AI. Rather than trying to control behavior, we show how four principles from contemplative traditions can help intrinsically align (super) intelligence. First, mindfulness enables self-monitoring and recalibration of emergent subgoals. Second, emptiness forestalls dogmatic goal fixation and relaxes rigid priors. Third, non-duality dissolves adversarial self–other boundaries. Fourth, boundless care motivates the universal reduction of suffering. We find that prompting AI to reflect on these principles improves performance on the AILuminate Benchmark ( d = .96) and boosts cooperation and joint-reward on the Iterated Prisoner’s Dilemma task ( d = 7 +). We also show how active inference offers parameters for integrating contemplative wisdom deeper into the architecture and world models of AI. This interdisciplinary approach offers a resilient alternative to brittle control schemes and may be the first empirical test of ‘ancient wisdom’.