What Stanford’s AI Transparency Scores Reveal

Stanford researchers assessed 10 foundation models across 100 indicators covering topics such as training data, human labor and computing resources. Llama 2 scored highest at 54 out of 100, while Amazon’s Titan scored 12; the report also drew criticism over how it measured transparency.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story focuses on limited disclosure and debate over measurement, without showing a clear shift toward AI danger or human decline.

What Stanford’s AI Transparency Scores Reveal

A Stanford report found major AI companies disclose limited information about how their foundation models are built. Its scores point to gaps in what businesses, researchers, policymakers and consumers can learn about the systems they may rely on—and its methodology has itself become part of the debate.

What the index measured

The Foundation Model Transparency Index assessed 10 popular models, including GPT-4. It used 100 indicators covering areas such as training data, labor practices and the computing resources used to develop each model.

The index treated disclosure as the measure for each indicator. In its data labor category, for instance, researchers asked whether companies explained which parts of the data pipeline involved human labor. The approach was intended to make visible information about both the inputs to models and the work behind them.

Foundation models are AI systems trained on large datasets that can perform tasks such as writing and generating images. Their role in generative AI expanded after the launch of OpenAI’s ChatGPT in November 2022. As organizations adopt these systems and fine-tune them for their own needs, information about limitations and biases becomes relevant to how they use them.

The scores showed room for disclosure

None of the models received a score the researchers considered impressive. Meta’s Llama 2 received the highest score, 54 out of 100. OpenAI’s GPT-4 scored 48, and Amazon’s Titan ranked lowest at 12.

Those results do not describe how capable a model is. They reflect how much information was disclosed against the index’s indicators. A score can therefore help readers compare reported transparency according to this framework, while leaving open questions about what should count as meaningful transparency.

That distinction matters to organizations considering whether to build applications on commercial models. If details about a model’s development and limits are hard to find, it can be harder to judge whether it is suitable for a particular use. The same lack of information can complicate academic research and make it more difficult for consumers to understand limitations or seek redress for harms.

Why openness matters beyond model builders

The researchers argued that disclosure affects more than companies choosing a model. Policymakers need information to develop meaningful policies for a powerful technology, while people affected by AI systems may need to know what those systems can and cannot do.

Stanford associate professor Dr. Percy Liang, who directs Stanford’s Center for Research on Foundation Models and advised on the paper, described a widening gap between openness and capability. He told Reuters, “It is clear over the last three years that transparency is on the decline while capability is going through the roof.”

The article described several pressures behind the move toward less open models, including competition among large technology companies and concerns about the dangers of spreading AI technology. OpenAI employees, in particular, had walked back the company’s previous open stance, citing potential dangers.

At TED AI, Liang also raised concerns about closed models such as GPT-3 and GPT-4 that do not provide code or weights. He discussed accountability, values and attribution of source material, and compared the possible benefits of open AI projects with collaborative efforts such as Wikipedia and Linux.

The transparency measure drew criticism

The index did not settle how transparency should be judged. After the article was published, AI experts, including Dr. Margaret Mitchell of Hugging Face, questioned its methodology. The AI ethics research lab EleutherAI wrote that the index’s claims were misleading about the spirit and facts of transparency in language models and could harm recent progress.

Neil Turkewitz, an AI critic and self-described “artists rights advocate,” made a related point in a reply to the author on X: “We care about transparency as a tool to achieve accountability, not for its own sake. As such, misleading or confusing metrics for transparency undermine their very purpose of promoting sensible governance through accountability.”

That challenge highlights a tension at the center of the report. Comparing disclosures can give organizations and governments a practical starting point, but the choice of indicators shapes what the scores say. The authors hoped the index would encourage companies to disclose more and help governments consider how to regulate the fast-growing AI field; critics argued that the measurement itself must support accountability clearly.