OpenAI Bets on AI to Help Solve Superintelligence Alignment

OpenAI is creating a Superalignment team led by Ilya Sutskever and Jan Leike to study how to control AI systems that could surpass human intelligence. Its plan is to use AI to help evaluate systems and eventually conduct alignment research, while acknowledging that the approach has risks and limits.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 1 ►

The story centers on controlling potentially superintelligent AI that could go rogue, though the effort is intended to improve safety.

OpenAI Bets on AI to Help Solve Superintelligence Alignment

OpenAI is setting up a team to tackle a difficult question: how could people steer AI systems that become more capable than humans? The company’s Superalignment team will research ways to keep such systems under control, using AI itself as part of the research process.

A team focused on control

The team will be led by Ilya Sutskever, OpenAI’s chief scientist and one of its co-founders, alongside Jan Leike, a lead on the company’s alignment team. Their work centers on the possibility that AI intelligence could exceed human intelligence within the decade.

That possibility makes control a central concern. Sutskever and Leike say there is currently no solution for reliably steering a potentially superintelligent system or preventing it from going rogue. They also point to a weakness in current alignment methods: techniques such as reinforcement learning from human feedback depend on people supervising AI, but people may not be able to reliably oversee systems much smarter than themselves.

OpenAI says the new team will work on the core technical challenges of superintelligence alignment over the next four years. It will include scientists and engineers from the company’s previous alignment division, as well as researchers from other parts of OpenAI.

Using AI to help with alignment research

The team’s proposed path is to build what its leaders call a “human-level automated alignment researcher.” In practical terms, the plan is to train AI systems with human feedback, use AI to assist in evaluating other AI systems, and eventually develop systems capable of doing alignment research.

Alignment research aims to make sure AI systems produce desired outcomes and do not go off the rails. OpenAI’s hypothesis is that AI could make progress on this research faster and more effectively than humans working alone.

The longer-term vision is for AI systems to take on a growing share of alignment work. Human researchers would then spend more effort reviewing research produced by AI systems, rather than generating all of it themselves. The company has described a future in which AI systems help develop improved alignment techniques and work with people to ensure their successors are more aligned with humans.

Risks and open questions

Having AI evaluate AI could expand the amount of work researchers can do, but it also creates a risk: flaws in an evaluator could spread through the process. The article notes that inconsistencies, biases, or vulnerabilities in an AI system could be scaled up when that system is used for evaluation.

There is also uncertainty about what makes alignment difficult. Some of the hardest parts of the problem may not be engineering challenges, which would limit what technical methods alone can accomplish. The proposal is a research direction, not a guarantee that a reliable control method will result.

Research intended to reach beyond OpenAI

Sutskever and Leike frame superintelligence alignment as a machine learning problem and say machine learning experts will be important to addressing it, including experts who do not already work on alignment. OpenAI also says it plans to share the results of the effort broadly.

The team’s stated scope includes contributing to the alignment and safety of models developed outside OpenAI. That makes the work relevant beyond the company itself, while leaving the central challenge unresolved: whether AI can help people supervise systems that may eventually outmatch them.