AI Tools Put Human-Curated Crowdsourced Data in Doubt

Researchers at Swiss university EPFL estimate that 33% to 46% of workers in one Mechanical Turk task may have used AI tools to help write summaries. The finding raises questions about whether crowdsourced data can still be treated as human-produced when tasks and worker incentives make AI assistance tempting.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

Possible AI-written summaries cast doubt on whether crowdsourced data still reflects human work and judgment.

AI Tools Put Human-Curated Crowdsourced Data in Doubt

A study of workers completing a writing task on Amazon’s Mechanical Turk suggests that AI assistance may be hard to spot in crowdsourced results. Researchers at Swiss university EPFL estimate that between 33% and 46% of workers in the task appeared to have “cheated” by using tools such as ChatGPT for some of the work. The finding points to a broader challenge: knowing whether data paid for as human-generated actually came from people.

Why crowdsourced work is used

Mechanical Turk connects people with tasks that computers still struggle to handle reliably. A requester submits work through an application programming interface, or API, and workers complete it before returning their results. Some tasks require human judgment because they are ambiguous, while the volume can be too large for a small team of experts.

One example is drawing bounding boxes to create datasets for computer vision models. The work can involve decisions that are not purely mechanical, yet may need to be done at a scale that would be difficult for a limited group of specialists. Crowdsourcing offers a way to distribute that work across many people.

But the value of this arrangement depends on what the requester expects workers to contribute. If a dataset is meant to capture human judgment or writing, results generated with an unknown AI tool may not serve the same purpose. The data may look complete while no longer representing the source its users expect.

A task suited to generative AI

To investigate the issue, the researchers asked Mechanical Turk workers to turn research abstracts from the New England Journal of Medicine into 100-word summaries. Summarizing text is also a task that generative AI tools such as ChatGPT are well suited to perform, which makes the exercise a useful test of whether workers might rely on them.

The researchers developed a method to distinguish human-written text from machine-generated text. They described that distinction as difficult for both people and machine learning models. Their estimate that 33% to 46% of workers appeared to have used AI applies to this particular task, rather than establishing how often AI assistance occurs across all crowdsourced work.

That uncertainty matters to organizations that use crowd workers to create or assess data. If AI-assisted work is mixed into a collection without being identified, a researcher may not know whether a sample reflects human responses, machine output, or a blend of both.

When the source of data matters

Datasets are treated differently depending on whether people or large language models produced them. A team evaluating its own language model against human work, for example, needs to know that its comparison set is actually human-generated. If the material instead comes from other models of unknown origin and quality, the evaluation may not answer the question the team set out to ask.

The same issue carries into training. The source article warns that training AI on machine-generated text can amplify bias or reinforce inaccurate information. When the provenance of text is unclear, people using the resulting dataset may have trouble judging what its patterns represent and how much confidence to place in results built from it.

The researchers argue that if workers use language models to complete crowdsourced tasks, the data loses value as a human gold standard. They also point to a practical mismatch: a requester might be able to prompt an LLM directly, potentially at lower cost, instead of paying workers to produce text that may have been generated with an LLM anyway.

Workers face incentives, too

There is a human side to the problem. The researchers note that crowd workers have financial incentives to use LLMs to increase productivity and income. For someone doing repetitive, low-paid work, an AI tool may seem like a way to finish tasks more efficiently.

Using tools to work faster is common in many jobs, but the arrangement becomes complicated when a requester is paying for a specific kind of contribution. A faster keyboard changes how a person types; an AI summarizer may change who or what produced the submitted text. Clear expectations about allowed assistance become important when the distinction affects how the results will be used.

The study does not settle how widespread AI use is across Mechanical Turk. It does show why dataset creators need to consider how they will identify the origin of submitted work. If human and machine-generated results cannot be reliably separated, the assumptions behind a dataset become harder to defend—and the trust placed in models trained or evaluated with it can weaken.