Spear phishing pulls names from LinkedIn and company sites to look personal; a BYU study found AI versions now match or beat human ones, and people told them apart only half the time.
AI-written spear phishing messages got clicked 28% of the time in a controlled BYU experiment, against 21% for human-written ones, and matched or outperformed human-written messages in roughly 80% of cases. Recipients could tell which was which only 52% of the time. The result is a market in which the surface cues readers have historically used to decide whether a message is legitimate can now be manufactured cheaply, by any attacker with a LinkedIn query and a language model.
Spear phishing, the targeted scam that pulls a target's job title, employer, and coworkers' names from LinkedIn and corporate websites, used to be expensive. The old ceiling on spear phishing throughput was the manual pretexting: finding a target's employer, looking up a coworker's name, drafting something plausible. AI collapses that ceiling to roughly the cost of a LinkedIn query, so any scammer can now produce a personalized message at volumes that used to be limited to bulk phishing. An AI that scrapes a target's recent conference talk, a coworker's social post about a project, and the company's open requisitions can draft a message that names all three. The target sees familiar context and clicks.
The new defense is verification through a channel the attacker does not control. If a message says it is from a coworker, the reader's job is to call that coworker on a number they already have, not reply to the message. If a message links to a vendor portal, the reader should navigate to the site directly rather than click the link. Security teams are accelerating the move to phishing-resistant authentication, including passkeys, hardware security keys, and platform-bound credentials: the password-reset email that pretends to come from IT is exactly the kind of surface cue AI can now synthesize at scale.
The BYU study, led by cybersecurity professor Derek Hansen with PhD student Jerson Francia, is one controlled experiment. It tested SMS messages, not email or voice, in a specific scenario design. The 28%-versus-21% click differential and the 2.3-times lift from mentioning a coworker are real for that setting, but real-world phishing performance is not measured here. The arXiv preprint and the KSL re-report both report the same headline numbers; the underlying mechanism, collapsing the per-target research cost, is what travels beyond the study.
Readers should stop asking "does this feel real?" and start asking "did I verify this on a channel I trust, independently of the message?" The first trusts the message's surface cues, which AI can now synthesize; the second trusts a separate channel the attacker doesn't control.