Why public AI usage data still leaves major blind spots

The AI Observatory analyzed real conversations with popular AI models and found that company reports can miss large parts of how people use generative AI. Its findings suggest personal, sensitive, and model-specific uses are harder to see when public evidence depends mainly on selective company reporting.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story is mainly about measurement gaps in AI usage reporting, with only mild concern around sensitive personal uses and transparency.

Why public AI usage data still leaves major blind spots

Public reports from AI companies offer a partial view of how people use tools such as Claude and ChatGPT. A new research project called the AI Observatory argues that the picture is still incomplete, especially when it comes to personal and sensitive conversations.

Why AI usage data matters

AI companies like Anthropic and OpenAI regularly publish reports about how people use their products. Those reports can shape how researchers, policymakers, and the public understand the benefits and risks of generative AI.

But researchers quoted in the source article say the companies decide what data to release. Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, said: “There is no independent source to corroborate it.”

Reuel is co-lead of the AI Observatory, a public platform built to aggregate and analyze real AI conversations. The project uses conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets.

The goal is not to replace company reports. It is to give researchers and policymakers an independent source of information when they assess how people are using generative AI.

What the AI Observatory found

The AI Observatory found that AI use differs significantly across models and has changed over time. Its research also shows more sensitive behavior than appears in reports from major AI companies, which the researchers say tend to focus more on work than on personal use.

Anthropic Economic Index is one of the best known sources of AI usage data. As its name suggests, it focuses on work- and productivity-related uses of Claude AI, filtering out conversations that are unrelated to those uses.

When the AI Observatory team applied Anthropic’s methods to its own dataset, nearly half of the conversations, or 48%, would have been filtered out. Those filtered conversations were more likely to include several sensitive or personal categories:

  • Health and relationships: 44.2% versus 31.2% in Anthropic’s analysis.
  • Adult or illicit topics: 7.9% versus 2.1%.
  • Harassment and hate: 27.5% versus 5.66%.
  • Sexual content: 16.7% versus 2.4%.

OpenAI’s 2025 report on ChatGPT use similarly found that only 30% of consumer use was related to work. That supports the broader point that work use is only part of the overall AI usage picture.

Different models, different behavior

The datasets studied by the AI Observatory include conversations from 2023 to 2025. Across that span, the researchers found differences in both user behavior and platform responses.

In WildChat, one of the largest and most detailed datasets included in the study, conversations became longer and more elaborate over time. The researchers saw increases in prompt tokens, response tokens, and conversation turns.

There was also significantly more small talk over time. The source article says that suggests AI companionship was increasing, while AI assistants’ self-disclosure that they were chatbots decreased.

At the same time, exchanges labeled as sensitive use dropped. The researchers used that label for conversations involving potentially harmful or restricted content, including sexual harassment and hate speech. That decline might suggest that platforms were deploying more effective safeguards.

The AI Observatory also found major differences depending on the model. Users varied in their topics, interaction styles, conversation structures, and in both the likelihood and type of sensitive use cases.

People used Grok and Gemini more frequently for information retrieval. Grok was especially popular for news and politics, but it was also where misinformation tended to concentrate. People were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.

The study also found differences among versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o.

Why company reports are not enough

Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel, said: “No single company report tells the whole story.” That is the central issue the AI Observatory is trying to address.

To create the platform, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions aggregated 24,521 conversations across 85,633 conversational turns. Those conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.

That dataset is still far smaller than what major AI labs can study internally. The latest Anthropic Economic AI Index is based on analysis of 1 million Claude conversations. OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.

There are also limits to the Observatory’s own data. Because the conversations come from voluntarily provided sources, the researchers caution that the findings are not indicative of all AI use. Sensitive uses may be underrepresented because people may be less likely to share them.

Even with those limits, the Observatory expands access for independent researchers. AI companies do not typically share their chat data for outside analysis, and independent researchers quoted in the source article say company reports can focus on findings that present the companies favorably.

The policy gap

The practical concern is decision-making. Stakeholders are making highly consequential decisions about AI’s benefits and risks with limited data, according to Reuel.

David Widder, an assistant professor at UT-Austin’s School of Information who researches how people interact with AI systems and is not involved with the AI Observatory, said the broader view helps researchers understand different uses more consistently. He noted that Anthropic has released separate blog posts about support or companionship and generating CSAM, but that a broader analysis is useful.

The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Reuel said the ideal scenario would be for AI companies to share their data with independent researchers in ways that protect user privacy.

Until then, the source article’s conclusion is straightforward: anyone relying on AI usage data is working with an incomplete picture. The available evidence shows that people use generative AI for work, but also for personal, social, sensitive, and model-specific purposes that public company reports may not fully capture.