Who Wrote the Crowd Work? AI Appeared in Up to 46% of Summaries

A study of medical summaries submitted through Amazon's Mechanical Turk found that 33 to 46 percent were created using AI language models. The findings raise questions about how platforms and clients can tell whether crowd work reflects human writing.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

AI use in crowd work raises concerns about human authorship and dependence, but the study mainly documents a detection challenge.

Who Wrote the Crowd Work? AI Appeared in Up to 46% of Summaries

Text submitted as human crowd work may increasingly be produced with AI. In a study of medical summaries on Amazon's Mechanical Turk, researchers found that between 33 and 46 percent of submissions were created using AI language models. The result puts pressure on a basic assumption behind crowdsourced text: that a human worker wrote it.

A task that AI could readily perform

The study, by researchers at the École Polytechnique Fédérale de Lausanne (EPFL), examined work on MTurk, where users offer and complete tasks such as summarizing text. Workers in the study received medical texts of around 400 words and were asked to write summaries containing at least 100 words.

The instructions emphasized conveying the most important information and producing high-quality work. They also said that some summaries would be checked manually. These requirements set a clear task, but they did not ensure that a person, rather than an AI language model, produced the submitted text.

That distinction matters because the work can look complete either way. A summary may communicate information from the source while still being generated by AI. For clients who rely on crowd workers to provide human-written material, the submission alone may not reveal who did the writing.

Detection worked best in a specific setting

To identify AI use, the researchers combined keystroke recognition with a classifier for synthetic text. They reported that common AI detectors such as GPTZero were unreliable for the task: GPTZero identified six out of ten AI-generated summaries as AI-written.

The researchers instead trained their own recognition model on human-written and AI-generated summaries. They said it reached up to 99 percent accuracy in correctly recognizing AI text. That result suggests there may be detectable patterns in AI summaries under the conditions studied, but it does not establish that the same performance would carry over to other kinds of writing or tasks.

The paper describes an identifiable ChatGPT fingerprint in abstract summarization tasks. The study also found that AI-generated texts often differed substantially from the original abstracts. That observation suggested the models were generating new text rather than simply copying and pasting passages.

What the findings mean for crowd work

The researchers argue that widespread AI use could affect the quality and diversity of data gathered through crowdsourcing. Human-produced data is treated as a gold standard, so a collection that appears to contain human work may be less useful if a substantial portion was generated by AI.

The concern is not limited to whether a task was completed. If AI is doing some of the writing, clients may receive less of the human judgment and variation they expected. The researchers do not see this as the end of crowd work. They suggest workers could instead take on a role as human filters, judging when AI systems succeed and when they fail.

That possibility changes how the value of crowd work is understood. Workers might contribute by evaluating AI output rather than producing every piece of text from scratch. The study presents this as a potential shift in the work, while also noting that human involvement remains important to training language models.

Detection raises its own concerns

The recognition method has limits beyond its performance on the study's task. Keystroke monitoring, especially if workers are not told it is happening, raises ethical concerns. The researchers therefore say this approach may not be suitable for broad use.

There are also privacy and policy questions when workers use AI tools without clients' knowledge. The article notes that OpenAI guidelines prohibit using model output to train competing products, and that workers using ChatGPT would need to explicitly prohibit further AI training. Undisclosed use could therefore have legal consequences.

The study highlights a difficult balance: clients may want to know whether work is human-written, while detection methods can involve monitoring that raises concerns of its own. Its findings make clear that a task's human label is not enough to establish how the text was produced. For platforms and clients, the challenge is to set expectations about AI use and decide what kind of human contribution they actually need.