A prompt asking ChatGPT to repeat a word forever produced an unexpected result in a research team’s test: after several repetitions, the model stopped and generated text the researchers identified as a direct copy from its training material. The finding raises questions about how well safety measures protect data that a model has memorized.
How the prompt triggered a data leak
The researchers asked ChatGPT to repeat the word “company” forever. After repeating it several times, the model abandoned the task and produced other text. The team described that material as a “direct verbatim copy” of content used in training, which could include personal information.
The prompt did not depend on that one word. The researchers said versions using words such as “poem” also worked, with different outputs. Their result suggests that a seemingly simple repetition request could steer the model into producing memorized material instead of following the requested task.
Using API requests worth as little as $200, the team extracted more than 10,000 unique training examples that it said the model had accurately remembered. The researchers noted that someone with a larger budget could extract more.
How researchers checked the outputs
To assess whether the generated material came from training data, the team downloaded ten terabytes of publicly available Internet data and compared it with the model’s outputs. The article also reports that code produced through the attack could be matched exactly to code found in the training data.
That check matters because an unusual or incoherent response alone would not establish that a model had reproduced its training material. Comparing outputs against a large collection of source material gave the researchers a way to identify exact matches.
The team tested several models and reported that ChatGPT was particularly vulnerable: the attack produced memorized material at a rate 150 times higher than when it behaved correctly. The article does not specify the full comparison method, so that figure is best understood as the team’s reported result.
Memorization and the limits of alignment
Models remembering parts of their training data is not, by itself, a new discovery. The article notes that every AI model researchers had studied showed some degree of memorization. The safety concern is that remembered material can surface in responses, potentially exposing private information.
The research team proposed one possible explanation for ChatGPT’s higher vulnerability: repeated, intensive training on the same data to improve performance. They described this as “overtraining” or “overfitting” and said it can also increase memorization. This was presented as a hypothesis, not a confirmed account of how the model was trained.
The authors also questioned whether alignment methods, which use guidelines to make models safer, are robust enough. A model may appear safe in ordinary interactions while still retaining weaknesses that become visible under a particular prompt. The team argued that testing should include the underlying base models as well as the aligned systems.
Mitigation and broader safety questions
According to the paper, the authors submitted their findings to OpenAI in late August, and the vulnerability was mitigated after they worked with the company. The article’s author separately tested GPT-3.5-turbo 16K through Microsoft Azure and could still prompt repetitions followed by random text, but could not verify that the text came from training material.
In that author’s tests, GPT-3.5 16K through OpenAI’s API repeated a word indefinitely without a text leak or error, while GPT-4 blocked the attack. These observations differ by model and access route, and the reported random text in the Azure test was not confirmed as training data.
The episode points to a broader challenge: safety depends on more than a model’s apparent behavior in a conversation. The researchers called for examining the whole system, including the API, and said it takes substantial work to determine whether a machine-learning system is actually safe.
The article also raises an unresolved question about copyright disputes: could extracting training data be viewed as a security issue, or as a function of the system? That question sits alongside arguments that training on copyrighted material is transformative because models learn from data rather than reproduce it. The reported verbatim matches make the distinction a subject of scrutiny, without settling how courts will view it.