AI short stories win readers over until authorship is revealed

Across three experiments with more than 2,500 total participants, readers struggled to tell ChatGPT-generated short stories from human-written ones. The AI stories received higher ratings for quality and immersion, but those judgments changed when readers were told who, or what, had written the text.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 2 ►

The story mildly points toward AI eroding human judgment and literary discernment rather than creating direct autonomy or safety risks.

AI short stories win readers over until authorship is revealed

Readers may be less able to separate human fiction from AI-generated fiction than they think. In a set of experiments involving short stories, participants could not reliably identify whether a text was written by a person or generated by ChatGPT.

The findings are not simply about detection. They also show how strongly expectations shape judgment: readers often rated stories differently depending on whether they believed the author was human or AI.

What the experiments tested

The research covered three experiments with more than 2,500 total participants. The central question was direct: can readers distinguish human-written fictional short stories from stories generated by ChatGPT?

In the first experiment, 1,682 participants each read one short story. There were six stories in total, each about 1,000 words long. Three were taken from well-known literary magazines and short story collections, while the other three were generated using ChatGPT 4.0.

The AI stories were not produced from generic instructions. Their prompts were based on the theme, style, and narrative perspective of the human originals, which made the comparison more focused on whether readers could perceive authorship from the finished work.

Participants were also given an authorship label. Half were told the story had been written by a human, and half were told it had come from ChatGPT. That information was accurate for only half the participants in each group.

Readers rated the AI stories higher

In the first experiment, the ChatGPT-generated stories were rated significantly higher than the human-written stories on perceived quality and immersion.

For quality, AI stories had a mean score of 1.54, while human stories scored 0.97 on a scale from minus 3 to plus 3. For immersion, the AI stories scored 1.42, compared with 1.00 for the human-written texts.

Those numbers matter because they show that readers were not merely failing to detect the AI stories. In this setup, they often responded to them more positively.

But the authorship label changed the picture. Regardless of who actually wrote the story, participants gave higher scores when they were told a human had written it. That suggests readers brought assumptions about AI writing into the evaluation itself.

Attitudes toward AI changed the response

The study also found that participants' own views of AI influenced their ratings. People with a positive attitude toward AI generally gave higher ratings across the board.

When those AI-positive participants were also told that a story came from ChatGPT, their scores increased further. Among AI-skeptical participants, the effect moved in the opposite direction.

This makes the result more complicated than a simple contest between human writers and AI systems. The same story could be judged through different filters depending on what a reader believed about the author and what they already thought about AI.

An earlier study on AI-generated poems found a similar bias, according to the source article. The broader pattern is that readers may react not only to the text, but also to the idea of who produced it.

Direct comparison did not solve the problem

Two additional experiments included 905 total participants and made the task harder. Instead of reading one story, each person read both a human-written story and an AI-generated story, then had to decide which was which.

Even with both texts in front of them, participants performed no better than chance. Self-reported experience with AI systems correlated positively with the ability to identify the stories' origins, while self-reported experience with fiction did not help participants tell them apart.

That distinction is important. Familiarity with fiction did not appear to give readers a reliable advantage in spotting AI-generated short stories. Experience with AI systems, however, was connected with better identification.

Why higher ratings are not the whole story

The researchers did not treat higher ratings as proof that AI writing was better in a literary sense. AI-generated texts tend to be smoother, easier to read, and more emotionally upbeat than human-written texts, according to the authors.

Those qualities may help explain why readers rated the AI stories highly. People tend to prefer material that is easier to process, so a smoother story can feel more satisfying even if that does not settle questions of literary depth.

High-quality literary fiction can work differently. It may be intentionally difficult, asking readers to spend more effort on meaning. The researchers suggest that a story can be high quality without being highly engaging, and a story can be engaging without being high quality.

The short story format may also favor AI. Producing a coherent story in 1,000 words is a different challenge from sustaining one across hundreds of pages. Within that shorter format, the study's conclusion is clear: AI can generate creative work that readers perceive as at least on par with human work, even while many people doubt that AI is capable of doing so.

All data and materials from the study are freely available on the Open Science Framework. The study was conducted by Sydney Sears and Deena Skolnick Weisberg and published in the journal Judgment and Decision Making.