arXiv · 2412.16468
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
Abstract
The emergence of large language models (LLMs) has sparked discussion on Artificial Superintelligence (ASI), a hypothetical AI system that surpasses human intelligence. Although ASI remains hypothetical and far beyond current AI capabilities, discussing its potential and exploring its feasibility and potential risks is critical for the development of future AI systems. The idea of superalignment originates from scalable oversight, which studies how to supervise increasingly capable AI systems when direct human supervision becomes insufficient. In this paper, we focus on the superalignment problem: "The process of supervising, controlling, and governing artificial superintelligence." We first review scalable oversight paradigms-Sandwiching, Self-Enhancement, and Weak-to-Strong Generalization -- then analyze the limitations of current paradigms through the lens of possibility and impossibility, discuss key challenges, and propose pathways for the safe and continual improvement of future AI systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
HyunJin Kim, DongHyun Ryu, Xiaoyuan Yi, Jing Yao, Jianxun Lian, Muhua Huang, Shitong Duan, JinYeong Bak, Xing Xie. 2024-12-21. The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment. https://arxiv.org/abs/2412.16468
Cite the original work for its findings. Save a collection to share your selection of sources.