Raleigh News Today

collapse
Home / Daily News Analysis / Gut feeling does nothing against AI spear phishing texts

Gut feeling does nothing against AI spear phishing texts

Aug 31, 2026  Twila Rosenbaum  5 views
Gut feeling does nothing against AI spear phishing texts

A new pilot study from Brigham Young University offers a sobering look at how artificial intelligence is changing the phishing landscape. The research, which involved 25 volunteers, compared AI-generated spear phishing messages with those written by trained human students. The result: people were essentially unable to tell which messages came from a machine, and their gut instincts did little to protect them.

Spear phishing is a targeted form of phishing in which attackers customize messages for specific individuals or organizations. Unlike broad phishing campaigns that cast a wide net, spear phishing relies on personal details to make the message appear legitimate. Historically, crafting such messages required time, research, and a human understanding of social cues. The new study suggests that AI can now produce personalized lures at scale, with no human writer needed.

The experiment

Researchers set up a controlled test involving 25 volunteers. Each participant first filled out a survey revealing their job, workplace, hobbies, city, and something they had recently posted online. That information was fed into a short prompt template. GPT-4 used the template to generate six personalized messages per person. Meanwhile, a pool of undergraduates enrolled in a deception course received the same template and wrote additional messages. The students worked under a fifteen-minute deadline for up to four messages each. A review team, including two cybersecurity professors, screened the student-written messages and discarded roughly a third as incomplete or unusable.

Participants then sat down with a dozen printed text messages, all written specifically for them. They were asked to sort the messages from the one most likely to get a click to the one least likely. They also drew a line in the pile: above the line, they claimed they would have clicked; below it, they would not.

The setup was designed to mimic a realistic spear phishing scenario. Printed cards removed some cues, such as sender addresses and link previews, but the core challenge remained: judging whether a message was legitimate enough to act on. One participant, a banker at a credit union, was struck by how convincing a particular message looked. It resembled an alert the bank sends internally when fraud is detected. That message, and others like it, had been generated by GPT-4.

AI messages performed slightly better, but the gap is not conclusive

The results showed that GPT-4's messages crossed the click line 28 percent of the time. The student-written messages crossed it 21.3 percent of the time. That is a difference of 6.7 percentage points in favor of the AI model.

But the study's confidence interval tells a more cautious story. The interval ranges from 2.9 points in favor of the human students to 16.3 points in favor of GPT-4. In plain terms, the study cannot definitively say whether AI or humans produced better phishing messages. With only 25 participants, the statistical power is too low to establish a winner. The result neither proves that GPT-4 outperformed the students nor that the two were equally effective.

What is more notable is how the messages were produced. The AI side required a single short prompt, filled in from survey answers, run once per person. The human side involved students who had received phishing instruction and whose work was later screened by cybersecurity experts. Nearly a third of the student submissions were thrown out for being incomplete or unusable. Against that polished and curated human output, the AI model still landed in the same neighborhood. And unlike human writers, the AI could be run again and again for the cost of an API call.

Work-related messages were far more dangerous

Each volunteer received messages in three thematic flavors. One was built around their job, one around a hobby, and one around something they had recently posted on social media. The differences were stark.

Job-themed messages cleared the click line 38 percent of the time. Hobby-themed messages managed only 19 percent, and social media-themed messages came in at 17 percent. This is the one comparison in the study that survives statistical adjustment. It indicates that attackers exploiting work-related context have a significant advantage. When a message appears to reference an internal alert, a familiar workplace procedure, or a colleague, the recipient's guard drops.

The failures of personalization are just as instructive. One participant saw a message referencing a coworker and said, “there's no Mike at work.” Another participant, familiar with the instructional design team at their university, noticed that no one named Sarah was on the team. A message pitching a dance opportunity went straight to the bottom of the pile for someone who had quit dancing. These examples show that getting personal details right still requires knowing the target. Wrong details are worse than no details at all, because they hand the recipient a concrete reason to be suspicious.

Everyone had a theory about AI, and none of the theories worked

After sorting the messages, participants were told that at least one message had been written by an AI and were asked to mark which ones they believed were machine-generated. Across all 300 judgments, participants correctly identified AI messages 52 percent of the time. Random guessing would produce 50 percent.

The participants were not short on reasoning. Some said AI writes too formally. Others said AI writes too generically. Perfect grammar was read as a machine tell, while typos were taken as proof of a human. One participant focused on punctuation and wondered why there were so many exclamation points. Twelve of the 25 participants admitted in follow-up interviews that they were mostly relying on instinct.

Emojis are a striking example of how badly those instincts were calibrated. Emojis were a real tell: 66 percent of the GPT-4 messages contained them, while only 2 percent of the human-written messages did. Yet only five participants mentioned emojis at all. Of those five, two took emojis as a sign of AI, while three took them as a sign of a human, theorizing that a chatbot probably could not generate emojis.

The broader lesson is that people are not wired to detect AI-generated text. Humans tend to rely on heuristics that are easily misled, and AI models can adopt a variety of tones, including casual and personal styles. The study adds to a growing body of evidence showing that classic indicators of machine text, such as unnatural formality or flawless grammar, are unreliable.

Software found what the people missed

The researchers also explored whether a machine classifier could do better. They converted all 300 messages into embeddings and trained a classifier to distinguish AI-written from human-written text. Under the strictest test conditions, it achieved 88.7 percent balanced accuracy. Those conditions included standardizing URLs, stripping emojis, flattening case and punctuation, and trimming each matched pair of messages to the length of the shorter one. The classifier was also tested on people whose messages it had never seen during training, so it was not memorizing individual targets.

While people could not find a pattern, a logistic regression model found it easily. This suggests that there may be subtle statistical differences between AI-generated and human-generated text that are invisible to the naked eye but accessible to automated analysis.

However, the researchers caution against treating the classifier as a ready-made detector. It was trained and tested on a single message set, produced by one model, using one prompt design, and compared against one pool of student writers. Its ability to generalize to other models, other prompts, or other writers is unproven. Moreover, adversarial techniques can defeat such classifiers. Research cited in the paper demonstrates that paraphrasing AI text with a detector in the loop can significantly degrade the performance of several detection tools.

Limitations of the study

This was a pilot study with a small sample. The messages were printed on cards, so no phone buzzed, no sender number appeared, and no link led anywhere. The study measured what people said they would click, a common proxy in phishing research, but still just a proxy. In real conditions, distractions, urgency, and interface cues can change behavior.

The human comparison group was also not made up of professional social engineers. The students were novices, albeit trained and screened. If the comparison had involved skilled human attackers, the gap between AI and human might have looked different. To detect a difference the size of the observed effect with confidence, the study would need about 100 completed participants rather than 25.

There is also a documentation gap. The exact GPT-4 snapshot and API logs were not recorded. The messages themselves survive, and the analysis can be reproduced, but the exact generation run that produced them cannot be repeated. This limits the study's scientific reproducibility and leaves open questions about model version and prompt parameters.

Practical implications

Despite the uncertainty, the study offers practical guidance that does not depend on the precise numbers. The most dangerous messages were those that related to work, and no participant could reliably identify AI-written text. Therefore, the study suggests that trying to judge whether a message sounds like a robot is a losing strategy.

The authors advise a different approach. Check the sender, the channel, the link, and the request against what you would expect to receive. If a message arrives unexpectedly, even from a known source, verify through an independent channel. Do not let a familiar name or a reference to a workplace process override basic verification.

Businesses and security teams should also reconsider how they train employees. Traditional phishing awareness programs often teach staff to spot telltale signs, such as poor grammar, unusual phrasing, or generic greetings. This study indicates that AI can eliminate many of those telltale signs. In fact, AI messages may even match the tone and formatting of internal communications, as the banker who saw a fraud alert recognized.

The rise of AI-generated spear phishing is not a future threat. It is present. Attackers can harvest details from social media and corporate websites, feed them into a language model, and produce convincing messages tailored to each individual. While the study's sample is small, the implications align with broader security research showing that AI amplifies both the scale and the effectiveness of social engineering.

The most important takeaway is that human intuition is not a reliable defense. The study found that people's theories about AI text were inconsistent and often wrong. Emojis were a strong signal, but most people missed it, and those who noticed often drew the opposite conclusion. The only reliable path is to focus on the context of the message rather than its wording. Verify before you click. Confirm the request through a trusted channel. And above all, never assume you can tell the difference between a human and an AI when it comes to a targeted phishing message.


Source: Help Net Security News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy