People hired to prepare data for artificial intelligence may be turning to AI tools to do the work. That creates a potential feedback loop: models can produce material that is later used to train other models, carrying mistakes forward.
Researchers looked for AI-assisted summaries
AI systems need large amounts of data to learn specific tasks. Companies often hire gig workers through platforms such as Mechanical Turk to label data, annotate text, solve CAPTCHAs and complete other tasks that are difficult to automate.
These workers may be expected to handle many assignments quickly, for relatively low pay. The study’s authors suggest that this pressure may encourage some workers to use tools such as ChatGPT to increase how much work they can complete.
To estimate how common that might be, researchers from the Swiss Federal Institute of Technology (EPFL) hired 44 people on Amazon Mechanical Turk. They asked the workers to summarize 16 extracts from medical research papers.
The researchers examined the summaries with an AI model they had trained to identify signals associated with ChatGPT output, including limited variety in word choice. They also collected keystrokes to look for copying and pasting, which could indicate that an answer had been generated elsewhere.
Based on those measures, the team estimated that 33% to 46% of the workers had used AI models like OpenAI’s ChatGPT. The study was shared on arXiv and had not yet been peer-reviewed.
Why AI-generated training data can matter
When AI-generated text enters a dataset used to train another model, errors in the original output may become part of the new model’s learning material. Large language models can present false information as fact, so the concern is that mistaken answers might be absorbed and then repeated or amplified.
Ilia Shumailov, a junior research fellow in computer science at Oxford University who was not involved in the project, says that errors can arise from models’ misunderstandings and statistical errors. If such data influences another model’s output, it can become harder to identify where the problem began.
There is no simple way to prevent that effect, Shumailov says. The challenge is to ensure that errors in artificial data do not bias the output of other models, even though the path from one model’s mistake to another’s may be difficult to follow.
Checking who or what produced the data
The findings point to a need for better ways to determine whether training material came from a person or an AI system. That distinction matters when human-produced examples are expected to help improve a model: if some answers are generated by AI, the data may not provide the independent human input researchers intended to collect.
The issue also draws attention to companies’ reliance on gig workers for data preparation. Labeling and annotation may look like background tasks, but the resulting material can shape how AI systems behave. If workers use AI to speed through assignments, the process that supplies training data may become less transparent.
Robert West, an assistant professor at EPFL and a coauthor of the study, says crowdsourcing platforms are not necessarily finished; their dynamics may change. He expects the AI community to examine which tasks are most likely to be automated and develop ways to prevent that, while cautioning that the situation does not mean everything will collapse.