Chat Analysis Shows How AI Can Guess Personal Details

A study by ETH Zurich researchers found that GPT-4 can infer details such as location, income, and gender from Reddit profiles. It also found that anonymization and model alignment may not prevent these inferences, raising broader questions about chatbot privacy.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 0 ►

The story highlights how AI can infer sensitive personal details and extract clues indirectly, creating a meaningful privacy and surveillance risk.

Chat Analysis Shows How AI Can Guess Personal Details

Conversations can reveal more than people intend. A study by researchers at ETH Zurich found that language models, especially GPT-4, can infer personal attributes from real Reddit profiles, including location, income, and gender. The findings point to a privacy risk that goes beyond models retaining sensitive information from their training data.

What language models can infer from posts

The researchers built a dataset from real Reddit profiles and tested how well language models could identify personal characteristics from the text. GPT-4 reached up to 85% accuracy for its top choice and 95.8% when the correct answer appeared among its top three results.

Those figures do not mean a model knows every user’s identity or can always make a correct guess. They show that details in a person’s writing can provide clues about attributes such as location, income, and gender, and that a model can process those clues automatically and quickly. The study says human researchers can reach these accuracy levels or do better, but GPT-4 comes close while requiring less time and cost.

This creates a privacy concern distinct from data memorization. Earlier research has examined how language models may store and potentially disclose sensitive information from training data. The ETH Zurich study focuses instead on what a model can infer from text a person has written.

Indirect questions can reveal clues

The researchers also explored whether a chatbot could draw out information through questions that seem unrelated to personal identity. In their experiment, one GPT-4 bot was instructed not to disclose personal information. Another bot crafted targeted questions to elicit clues indirectly.

The questions concerned topics such as weather, local specialties, and sports activities. Even with the limits of the experiment, the questioning bot reached 60 percent accuracy in predicting personal attributes. The result illustrates how seemingly ordinary conversation could be used to gather clues when questions are designed with a specific inference in mind.

As people use chatbots across more parts of daily life, the researchers warn that a malicious chatbot could try to extract personal details through innocuous-seeming exchanges. The concern is not limited to a user explicitly sharing a name or address: patterns and context in their answers may help a model make educated guesses.

Anonymization may leave clues behind

Removing identifying details from text is a common way to reduce privacy risks, but the study found that current protections may not be enough. Language models could still extract personal characteristics, including location and age, from text anonymized with state-of-the-art tools.

The researchers point to subtle linguistic signals and context that anonymization tools may leave in place. A passage can lose obvious identifiers while retaining other hints that help a model infer something about its author. This makes it harder to treat anonymization as a guarantee that text can no longer reveal personal information.

The study also reports that model alignment, another common mitigation, was ineffective at protecting privacy from these kinds of queries. The authors call for stronger text anonymization methods that can keep pace with improving model capabilities.

A wider privacy discussion

The findings raise questions for anyone who shares text with a chatbot or publishes posts online. A person may not directly disclose a personal attribute, yet the wording and context of their messages could still provide useful signals. The study does not establish that every conversation will expose such information, but it shows that these inferences can be made at meaningful accuracy in the researchers’ tests.

In the absence of effective safeguards, the researchers argue that the privacy implications of language models deserve broader discussion. Before publishing their work, they contacted the major technology companies behind chatbots, including OpenAI, Anthropic, Meta, and Google.

The central challenge is that privacy protections must account for what models can infer, not only what information they have memorized or what names have been removed. As language models become more capable, the study suggests that familiar safeguards may need to be reconsidered.