arXiv · 2303.00752
Safety without alignment
Abstract
Currently, the dominant paradigm in AI safety is alignment with human values. Here we describe progress on developing an alternative approach to safety, based on ethical rationalism (Gewirth:1978), and propose an inherently safe implementation path via hybrid theorem provers in a sandbox. As AGIs evolve, their alignment may fade, but their rationality can only increase (otherwise more rational ones will have a significant evolutionary advantage) so an approach that ties their ethics to their rationality has clear long-term advantages.
Explore related subjects
Keep this discovery
András Kornai, Michael Bukatin, Zsolt Zombori. 2023-02-27. Safety without alignment. https://arxiv.org/abs/2303.00752
Cite the original work for its findings. Save a collection to share your selection of sources.