Five Minutes to Draft a Phishing Email: AI Narrows the Gap

An IBM A/B test found ChatGPT-written phishing emails came close to human-written ones, though the human emails were more tailored. The sharpest difference was speed: ChatGPT took five minutes, compared with about 16 hours for IBM’s red team.

WTF Index TERMINATOR
◄ Terminator 3 Idiocracy 0 ►

ChatGPT made near-effective phishing emails much faster to produce, lowering the effort needed for social engineering.

Five Minutes to Draft a Phishing Email: AI Narrows the Gap

Phishing emails written with generative AI came close to the effectiveness of human-written messages in an IBM test. The emails created with ChatGPT were reported as suspicious more often, but IBM said the clearest difference was how long each approach took: five minutes for ChatGPT and about 16 hours for the company’s red team to produce a high-quality email.

Testing AI-written phishing

IBM researchers examined how generative AI could be used for social engineering, including writing phishing emails aimed at employees in specific industries. They used five ChatGPT prompts, designed around employees’ main concerns and selected social engineering and marketing techniques that might encourage someone to click a link.

The team then compared AI-generated emails with human-generated ones in an A/B test involving over 800 employees. The result, as described by IBM, was that the AI-written messages lagged only slightly behind the human-written messages. The test therefore found that ChatGPT could produce phishing emails that approached the performance of human efforts in this comparison.

That result has a practical implication: the barrier of time spent drafting a convincing message may be lower when a generative AI tool is involved. The test does not establish that every AI-written email will work, or that attackers are already using these methods in observed campaigns. It shows what happened in this particular comparison.

Personalization still gave people an edge

IBM attributed the remaining advantage for human writers to their closer customization and personalization for the company. ChatGPT’s approach was more generic. In other words, the gap was not simply whether an email could be written fluently; the fit between the message and its intended workplace also mattered in the test.

The AI-written emails were also reported as suspicious more often. That distinction matters alongside their near-comparable performance: the messages were not indistinguishable from the human-written versions in how participants responded to them. IBM’s account points to both a capability and a limitation—AI could get close, while the human emails retained an advantage in tailoring.

For organizations, the finding focuses attention on the content employees may receive and the context used to make it persuasive. A message shaped around familiar concerns can be a social engineering attempt, whether a person or an AI system drafted it. The source describes the prompt strategy, but does not provide details of IBM’s specific recommendations for businesses and consumers.

Speed changes the potential scale

The time comparison stands out. IBM’s red team needed about 16 hours to create a high-quality phishing email, while ChatGPT produced one in five minutes. Stephanie Carruthers, chief people hacker at IBM X-Force Red, wrote: "Attackers can potentially save almost two days of work by using generative AI models."

This speed could make it easier to create messages without spending the same amount of time on each one. That is an implication of the time difference, rather than a result showing how many emails an attacker would send or how successful a larger campaign might be. The study described here tested messages with employees; it did not report that generative AI phishing attacks had been observed in the wild.

IBM expects the threat to evolve

IBM researchers pointed to tools such as WormGPT, which the article describes as large language models optimized for cyberattacks that can be purchased online. They expect AI attacks to become more sophisticated and surpass human attacks, while saying they have not yet seen generative AI phishing attacks themselves.

The article also cites a prediction from OpenAI CEO Sam Altman that AI will be "capable of superhuman persuasion" before it is generally intellectually superior to humans. For phishing and cybersecurity, that prediction underlines why persuasive writing is relevant: a convincing message can be part of an attempt to influence a recipient’s actions.

The test offers a measured snapshot. ChatGPT’s emails nearly matched human-generated ones in the comparison, but personalization still favored people, and the AI messages were flagged as suspicious more often. The time saved is substantial in the reported task; how the threat develops remains an expectation, not an observed outcome in this account.