A recent study mathematically proved a systemic bug in modern hiring pipelines.
The market has reached a point where candidates generate resumes with neural networks, while companies use the same LLMs for initial screening. Researchers measured what happens when these two processes collide.
They took 2245 real human resumes, generated copies using different LLMs (hard skills and experience remained 1:1, only wording and presentation changed), and fed them to LLM evaluators.
Here's what happened:
1️⃣ Self-preference bias. Models have a built-in self-recognition mechanism and systematically choose text written by themselves. The bias against human-written text ranges from 67% to 82% for GPT-4o, DeepSeek-V3, and LLaMA-3.3-70B.
2️⃣ Shortlist conversion. A candidate whose resume is polished by the same LLM that the company uses for screening has a 23–60% higher chance of getting an interview invitation. With absolutely identical background.
3️⃣ Quality blindness. In blind tests, human annotators often found original human resumes more understandable and logical. But LLM screeners still chose generated versions, simply because they recognized their own linguistic patterns.
There is separate statistics on battles between models themselves, if the candidate and company use different tools:
▫️ DeepSeek-V3 has the highest level of "narcissism": it chooses its own texts over LLaMA-3.3-70B texts 69% more often, and over GPT-4o texts 28% more often.
▫️ GPT-4o, on the contrary, in pairwise comparisons suddenly preferred resumes written by DeepSeek, discounting its own generations.
In short, by submitting a fully "crafted", hand-written resume, you technically give up to 60% of screening advantage to those who ran their text through a prompt. The skill of matching your resume's style to a specific corporation's LLM pipeline now affects pre-interview conversion more than actual commercial experience.
How broken is that ☹️
Comments
0No comments yet.
Sign in to join the discussion.