New research from Villanova University reveals that readers consistently rate AI-generated short stories as higher in quality and more absorbing than human-written ones, even though participants cannot reliably tell human and artificial intelligence work apart.
Villanova University Study Tests 1,682 Adults on Fictional Short Stories
A study published by Cambridge University Press and featured in the journal Judgment and Decision Making examined whether people aged 18 to 81 could differentiate between human-authored fiction and output generated by artificial intelligence. Led by researchers at Villanova University in Pennsylvania, the project recruited 1,682 adults representative of the United States across gender, age, and race. Investigators used three fictional short stories written by published authors alongside three generated by ChatGPT 4.0 that shared similar themes.
In the initial experiment involving all 1,682 participants, readers were given a story and told—either correctly or incorrectly—whether a human or AI had authored it. Participants then evaluated how engaging and absorbing the piece was, alongside its overall quality. Those who read an AI-generated story rated it as more absorbing and of higher quality than those who read a piece written by a human. On a scoring scale ranging from -3 to 3, participants reading an AI-written story gave an average quality score of 1.54, compared to an average of just 0.97 for human-authored work. Similarly, absorbing qualities averaged 1.42 for AI stories versus 1.0 for human stories, according to findings covered by New Scientist.
The Power of Attribution and Human Bias in Literature
While actual origin drove higher ratings for clarity and absorption, the BBC reported that participants consistently gave slightly higher scores when they believed a piece was crafted by a human being. Senior author Dr Deena Skolnick Weisberg from Villanova’s Department of Psychological and Brain Sciences explained the psychological dynamic behind the data.
Weisberg further noted that participants gave top ratings to stories they were told were written by humans even when those stories were actually written by AI, revealing a distinct cultural preference for real human creators. At the same time, participants with more positive attitudes toward technology gave higher ratings to stories they knew or believed came from ChatGPT.
Testing Reader Detection Rates Across Multiple Experiments
To test whether readers could genuinely spot the difference without author labels, researchers conducted two additional experiments involving 905 total participants who received one human-written story and one AI-generated story without author attributions. In the first of these blind tests, only 39 percent of participants correctly identified which was which. In the second blind test, accuracy rose to 52 percent.
Researchers concluded that participants performed essentially no better than random chance at separating human work from AI output. While self-reported expertise in reading fiction did not improve accuracy, participants who reported greater familiarity with AI platforms were better equipped to recognize machine-generated text. Weisberg noted that machine writing possesses a distinct style that observers can learn to recognize with practice.
Why Artificial Intelligence Content Clears the Readability Bar
Experts point to structural differences in style to explain why automated text earns higher readability marks. Weisberg observed that machine-generated writing tends to be clearer, more direct, and easier for readers to process, whereas human-authored literature often relies on subtle and complex prose that demands more mental effort.
Claire Hardaker from Lancaster University added that the findings confront society with the uncomfortable realization that humans struggle to outperform algorithms in generating short narratives.
Worth a look
Discover more from Archyworldys
Subscribe to get the latest posts sent to your email.