Why ChatGPT’s training data is raising privacy questions

Data regulators are investigating how OpenAI collected and uses information in ChatGPT, including personal details and chat logs. The case highlights a practical challenge: people may have rights to access, correct, or erase data, while identifying that data inside large AI training sets can be difficult.

WTF Index TERMINATOR
◄ Terminator 3 Idiocracy 1 ►

The story focuses on privacy risks from AI data collection and chat logs, raising concerns about surveillance and control of personal information.

Why ChatGPT’s training data is raising privacy questions

OpenAI’s approach to gathering data for its AI systems is facing scrutiny from privacy regulators. Investigations in several countries, along with Italy’s temporary block on ChatGPT, have put a difficult question into focus: how can a company meet people’s data rights when information has been gathered from the internet and used to train models?

Regulators are asking how ChatGPT uses personal data

Data protection authorities in France, Germany, Ireland, and Canada are investigating how OpenAI collects and uses data. The European Data Protection Board is also setting up an EU-wide task force to coordinate investigations and enforcement around ChatGPT.

Italy has taken a precautionary step by blocking ChatGPT’s use and giving OpenAI until April 30 to comply with the law. The company would need to obtain consent for collecting people’s data or show that it has a “legitimate interest” in doing so. It would also need to explain how ChatGPT uses the data and give people ways to correct errors, request erasure, and object to its use.

The concerns extend beyond information used to train a model. Regulators have also questioned how OpenAI handles data gathered after training, including chat logs. People may share sensitive details about their health, mental state, or personal opinions in conversations with a chatbot. If such information can be repeated to others, that creates a privacy concern.

Publicly available information can still carry rights

OpenAI says its models are trained on publicly available content, licensed content, and material generated by human reviewers. The company has said it believes it complies with privacy laws and works to remove personal information from training data upon request “where feasible.”

But public availability alone does not settle the question under European data protection rules. People can have rights as “data subjects” even when information about them was public. Those rights include being informed about how data is collected and used, and asking for it to be removed from systems.

That leaves OpenAI with a legal question about its basis for collecting data. If it cannot show that people consented, it may need to make the case that collecting the information serves a “legitimate interest.” That argument would require regulators to weigh the company’s purpose against people’s rights.

Removing data from a model is a difficult task

Meeting a request to erase information can be more complicated when the data has been gathered into a large training set. The source article describes how AI developers commonly collect material from the web, then process it to remove duplicates or irrelevant items, filter unwanted content, and correct errors. When records of those steps are limited, it may be hard to know exactly what information went into a model.

That creates a practical problem for identifying and removing a particular person’s data. Even if a company finds the relevant information in its training materials, it may be unclear whether removing it there would also remove its influence from a model. Copies of online data can also remain available after the original has been deleted.

The issue is not limited to training data. Chat logs raise a separate set of questions because people may give a service information directly in conversation. Regulators have asked whether users can have those logs deleted and how the service prevents sensitive details from being repeated.

A broader test for AI data practices

The outcome could shape how AI companies collect data, since regulators beyond Europe are watching the case. European rules have been copied widely, and the questions raised in Italy may have implications for companies that rely on large collections of online material.

Possible consequences for OpenAI include country-level or EU-wide restrictions, fines, or orders to delete data and models. The source article notes that resolving the legal questions could take years, potentially bringing the dispute before the Court of Justice of the European Union.

At the heart of the dispute is a gap between the scale of AI training and the ability to track individual pieces of information. The case puts pressure on companies to explain what they collect, establish a legal basis for using it, and provide a workable response when people ask to exercise their data rights.