Italy’s ChatGPT Order Puts Personal Data Practices Under Scrutiny

Italy’s data protection authority ordered OpenAI to stop processing people’s data locally while it investigates possible GDPR violations. The concerns include the legal basis for training, inaccurate information about people, a reported data breach and protections for minors.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story focuses on privacy risks from large scale data collection and processing, with a mild lean toward AI-enabled surveillance and control.

Italy’s ChatGPT Order Puts Personal Data Practices Under Scrutiny

Italy’s data protection authority ordered OpenAI to stop processing people’s data locally with immediate effect, citing concerns about ChatGPT and European privacy rules. The order opened an investigation into whether the service’s data practices comply with the General Data Protection Regulation (GDPR), and raised questions that reach beyond one chatbot.

What Italy is investigating

The authority, known as the Garante, said it was concerned that OpenAI had unlawfully processed people’s personal data. It also pointed to the lack of a system to prevent minors from accessing the service. OpenAI was given 20 days to respond to the order.

Under the GDPR, the rules apply when the personal data of people in the European Union is processed. ChatGPT can produce biographies of named individuals, while OpenAI has disclosed that earlier models were trained on data scraped from the Internet, including forums such as Reddit. The source article says OpenAI did not provide details of the training data for GPT-4.

The authority also highlighted that ChatGPT’s information about people does not always match real data. When a system produces inaccurate claims about a named person, that raises a practical question: how can the person ask for the information to be corrected? The GDPR gives Europeans rights over their data, including the right to rectification, but the article says it was unclear how or whether users could obtain that remedy from OpenAI.

Questions about lawful use and data security

A central issue is the legal basis for using people’s data to train large language models. The GDPR allows several possible legal bases, including consent and public interest. The Garante questioned whether the mass collection and storage of personal data for training had a valid basis, and raised concerns about transparency, fairness and data minimization.

These questions matter because training involves using information at scale, while individuals may not know their data has been repurposed. If authorities determine that data was processed unlawfully, they could order it deleted. Whether that would require retraining models built using that data is an unresolved question in the article.

The investigation also followed a data breach disclosed by OpenAI earlier that month. The company said a conversation history feature had exposed some users’ chats and may have exposed some users’ payment information. The GDPR regulates data breaches and requires companies to notify supervisory authorities of significant breaches within tight time periods.

Protecting minors is another concern

The Garante said it was concerned that OpenAI was not actively preventing people under the age of 13 from signing up, for example through age verification technology. Without a reliable way to confirm users’ ages, the authority could require accounts it cannot confirm do not belong to children to be deleted, followed by a more robust sign-up process.

The regulator had also acted over child safety concerns involving the virtual friendship AI chatbot Replika and pursued TikTok over underage usage. The source article describes those actions as context for Italy’s attention to children’s data, while the ChatGPT order specifically focuses on access controls and the possible processing of minors’ information.

A wider test for AI and privacy rules

Italy’s order could have implications beyond OpenAI. The article notes that OpenAI had no legal entity established in the European Union, and that under the GDPR a data protection authority can intervene when it sees risks to local users. Other authorities could therefore take an interest too.

Data protection expert Lilian Edwards told TechCrunch that the more fundamental question was the lawful basis for processing, which she said could affect many machine learning systems. She also noted that large language models do not offer the same clear remedies for correcting or removing information as search engines do. Enforced retraining was raised as one possible fix, though its consequences remain uncertain.

The order puts a difficult issue into focus: privacy rules give people rights over their personal data, while AI systems can be trained on large collections of information and generate claims about individuals. Italy’s investigation will examine whether OpenAI’s practices meet those rules and what remedies may be available if they do not.