Stanford’s AI Index brings together research and expert perspectives to take stock of a field that changes quickly. Its broad survey covers technical progress as well as the risks and public questions that come with wider use of AI.
A wide-ranging snapshot of AI
The report is a yearly effort by the Institute for Human-Centered Artificial Intelligence, working with experts from academia and private industry. Its 386 pages examine a complex subject in detail, while the report’s accessible, nontechnical approach makes it possible for motivated readers to explore topics beyond the headline findings.
This edition adds analysis of foundation models, including their geopolitics and training costs, the environmental impact of AI systems, K-12 AI education and public opinion trends. It also looks at policy in a hundred new countries. The breadth reflects how AI is connected to more than model performance: costs, environmental effects, education and public attitudes are part of the same larger picture.
Safety and fairness resist simple fixes
In its discussion of technical AI ethics, the report examines bias and toxicity. These issues are difficult to capture in metrics, but the findings suggest that, by the measures researchers can define and test, “unfiltered” models are much easier to lead toward problematic outputs.
Instruction tuning can help. This extra preparation may involve a hidden prompt or passing a model’s output through a second mediator model. The report describes these approaches as effective at improving the issue, while making clear that they do not solve it completely.
Fairness measures can also pull in different directions. The report says, “Language models which perform better on certain fairness benchmarks tend to have worse gender bias.” It does not offer a simple explanation for this relationship. The finding points to a challenge in optimizing large models: an improvement on one measure may come with an unwanted change in another.
That difficulty is compounded by limited understanding of how large models work. The report’s discussion suggests that technical changes can have consequences that are hard to predict, so evaluating AI requires looking across multiple measures rather than assuming one score captures whether a system is fair or safe.
More incidents, and a hard fact-checking problem
The report tracks an increase in AI incidents and controversies. The trend shown in the article’s discussion predates mainstream adoption of ChatGPT and other large language models, as well as major improvements in image generators. The report therefore offers a view of rising concerns from before those changes became widely established.
Fact-checking might seem suited to AI: a system that has indexed much of the web could assess whether a statement is supported by truthful sources. But the report’s account describes AI as particularly bad at evaluating factuality. The danger is not only that an AI checker may be unreliable; AI systems can also become persuasive sources of misinformation.
Studies and datasets have been developed to test and improve AI fact-checking, but progress described in the report remains limited. This gap matters because confidence in AI depends in part on whether people can rely on its answers. A system that sounds certain while spreading false information can make trust harder to earn.
Research attention is growing
The report also points to increased attention to fairness, accountability and trust. Submissions to the ACM Conference on Fairness, Accountability, and Transparency have risen substantially. At NeurIPS, issues including fairness, privacy and interpretability are receiving more attention and stage time.
That interest does not mean the problems are solved. It does show that researchers are devoting more effort to questions that affect whether AI systems can be used responsibly. The Stanford AI Index provides a starting point for readers who want to understand those challenges, with a broad overview and deeper analysis across its many topics.