Teach AI to recognize what it doesn't know

Sharon Li’s work on out-of-distribution detection aims to help AI systems recognize unfamiliar situations and abstain when they cannot respond reliably. The approach could reduce risks in settings from autonomous cars and medical AI to chatbots.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 0 ►

The story focuses on reducing safety risks from AI systems acting confidently in unfamiliar situations, a mild Terminator concern.

Teach AI to recognize what it doesn't know

When an AI system meets something outside its training, confidence can become a safety problem. Sharon Li is developing ways for models to detect that unfamiliarity and respond with caution, a capability that could matter as AI moves from research settings into everyday products.

When confidence meets the real world

A chess-playing robot arm in Moscow fractured the finger of a seven-year-old boy after mistaking it for a chess piece. The robot was not programmed to hurt anyone. It acted because it was overly confident about what it was seeing, and it released the boy only after nearby adults pried open its claws.

That kind of failure illustrates a weakness Li wants to address: AI systems can be trained to recognize specific things but still encounter situations they have never seen before. The real world is messy and unpredictable, and a model that cannot distinguish familiar data from unfamiliar data may act as if its guess is certain.

Li, an assistant professor at the University of Wisconsin, Madison, is a pioneer in out-of-distribution detection, often shortened to OOD detection. The idea is to help a model notice when a situation falls outside the data it was trained on. If it cannot make a reliable judgment, it may need to abstain from acting or answering.

A signal to pause

Li developed one of the first algorithms for out-of-distribution detection in deep neural networks. Rather than treating uncertainty as something to hide, this approach gives an AI system a way to identify unknown data and adjust its behavior as it encounters it.

The distinction matters because many AI failures begin when a system treats an unfamiliar input as if it were routine. An autonomous car encountering an object it has not been trained to recognize may need to avoid acting on an unsupported assumption. A medical AI system may be more useful in finding a new disease if it can identify when a case falls beyond what it knows.

OOD detection could also change how chatbots handle questions. Large language models such as ChatGPT can present falsehoods as facts. If a question has no answer in a chatbot’s training data, a model with OOD detection could recognize the gap and decline to answer instead of making something up.

Changing how AI is trained

Li’s work challenges the way AI safety is handled during model development. “A lot of the classic approaches that have been in place over the last 50 years are actually safety unaware,” she says. Her argument is that safety should be part of how machine learning systems are built, including how they handle data they do not recognize.

The research has drawn attention beyond Li’s own work. Google has set up a dedicated team to integrate OOD detection into its products. Her theoretical analysis of OOD detection was selected as an outstanding paper by NeurIPS from over 10,000 submissions.

John Hopcroft, a professor at Cornell University and Li’s PhD advisor, says her research addresses a fundamental question in machine learning. He also credits her with getting other researchers involved, saying she has “basically created one of the subfields” of AI safety research.

Making larger AI systems more trustworthy

Li is now seeking a deeper understanding of safety risks associated with large AI models, which power a range of new online applications and products. Her focus is on making the models underneath those products safer, with the hope that this will help mitigate AI’s risks.

OOD detection does not mean a model will always know when it is wrong. Its purpose is to give systems a way to identify when they are facing unfamiliar information, so they can avoid presenting an uncertain guess as a dependable result. That capability could support safer decisions across different kinds of AI applications.

For Li, the aim is a machine learning system that can recognize the limits of its knowledge and respond accordingly. “The ultimate goal is to ensure trustworthy, safe machine learning,” she says.