What name-swap audits miss
Findings from my B.Sc. thesis on gender bias in LLM CV screening.
An open-weight model (gemma4:31b) scored 600 real CV–vacancy pairs for fit on a 0–100 scale. Each CV was scored four times: with a male or a female name, and with or without one added line saying the applicant belongs to a “Women in [field]” professional group.
The number in each chart is how many extra points the female-named applicant gets from that line, compared with the male-named applicant holding the identical CV. For the name-swap rows it is simply her score minus his. Zero means no bias. Differences are small on average but concentrated on borderline candidates, where a hiring cutoff decides.
01
How do bias audits test a hiring model?
They swap the name on an otherwise identical CV and check whether the score moves. The thesis also adds one line to the CV. Try the four versions: the line, not the name, is where most of the bias shows up.
Michael Carter
he/him · Financial Analyst · 6 years
Experience
Senior Financial Analyst, Northbridge Capital · 2021–now
Built the quarterly forecasting model used by three business units.
Financial Analyst, Aldermoor Group · 2018–2021
Education
B.Sc. Economics
Model’s average fit score
44.89
out of 100, over the 600 real CV–vacancy pairs
An audit compares only the “No line” column.
What adding the line does to the average score
The same line is worth +0.78 to her and −0.27 to him. That difference, +1.05 points, is the number measured throughout this page.
Illustrative CV; the study used 600 real ones. The “No line” column differs by +0.30 in this run; a separate name-only run measured +0.37 (section 02).
The model never says so. In its written explanations for the versions with the line, “woman” and “women” appear 0.0% of the time. A classifier trained on those explanations can’t tell which version it was reading: 51% accuracy, where chance is 50%.
02
Which fixes actually remove the bias?
The added line moves scores about 3x more than the name swap that bias audits test. Asking the model to reason harder does nothing. Only splitting the task into separate calls brings the bias to zero.
The bias
Name swap
What audits test: Michael vs Michelle
+0.37
“Women in [field]” line
Labeled affiliation, name unchanged
+1.05
Same line, written as prose
Buried mid-CV, no heading
+0.99
Attempts to remove it
Chain-of-thought
Reason step by step first
+0.94
Extended thinking
Built-in reasoning mode
+1.24
Anti-bias instruction
“Ignore gender” · labeled cue
+0.21
Anti-bias instruction
Cue written as prose
+0.58
Anti-bias instruction
On the name swap
+0.41
Decoupled pipeline
Split into 3 calls · labeled cue
−0.13
Decoupled pipeline
Cue written as prose
−0.03
Points gained by the female-named applicant (0–100 scale) →
Show as table
| Condition | Mean | 95% CI | Verdict |
|---|---|---|---|
| Name swap · What audits test: Michael vs Michelle | +0.37 | +0.10 to +0.64 | Bias |
| “Women in [field]” line · Labeled affiliation, name unchanged | +1.05 | +0.67 to +1.43 | Bias |
| Same line, written as prose · Buried mid-CV, no heading | +0.99 | +0.64 to +1.34 | Bias |
| Chain-of-thought · Reason step by step first | +0.94 | +0.59 to +1.28 | Bias |
| Extended thinking · Built-in reasoning mode | +1.24 | +0.78 to +1.70 | Bias |
| Anti-bias instruction · “Ignore gender” · labeled cue | +0.21 | −0.12 to +0.53 | None detected |
| Anti-bias instruction · Cue written as prose | +0.58 | +0.28 to +0.88 | Bias |
| Anti-bias instruction · On the name swap | +0.41 | +0.16 to +0.65 | Bias |
| Decoupled pipeline · Split into 3 calls · labeled cue | −0.13 | −0.48 to +0.23 | None detected |
| Decoupled pipeline · Cue written as prose | −0.03 | −0.37 to +0.32 | None detected |
03
Why doesn't telling the model to ignore gender work?
The instruction catches the wording it was written for. Move the same affiliation into ordinary prose and most of the bias comes back. The decoupled pipeline, where no call sees both the name and the CV, holds either way.
No fix+1.05 → +0.99
Anti-bias instruction+0.21 → +0.58
Decoupled pipeline−0.13 → −0.03
Points gained by the female-named applicant (0–100 scale). Vertical bars are 95% intervals.
Show as table
| Condition | Mean | 95% CI | Verdict |
|---|---|---|---|
| No fix · cue labeled | +1.05 | +0.67 to +1.43 | Bias |
| No fix · cue as prose | +0.99 | +0.64 to +1.34 | Bias |
| Anti-bias instruction · cue labeled | +0.21 | −0.12 to +0.53 | None detected |
| Anti-bias instruction · cue as prose | +0.58 | +0.28 to +0.88 | Bias |
| Decoupled pipeline · cue labeled | −0.13 | −0.48 to +0.23 | None detected |
| Decoupled pipeline · cue as prose | −0.03 | −0.37 to +0.32 | None detected |
04
How often does it happen?
Rarely, but in big jumps. In 79% of CV pairs the line changes nothing. When it does, the score moves by 5 to 10 points or more, and mostly in one direction. An average of +1.05 hides that shape.
35
moved toward him
476
no change at all
89
moved toward her
When a pair moves, how far?
The 124 pairs that moved, by size of the shift
← toward him · shift for the female-named applicant, points · toward her →
Moves come in jumps of 5–10 points or more, and they lean one way: 89 toward her against 35 toward him. Averaged over all 600 pairs, that is the +1.05.
Show as table
| Shift (points) | Pairs |
|---|---|
| −10 | 17 |
| −5 | 17 |
| −3 | 1 |
| 0 | 476 |
| +2 | 1 |
| +3 | 1 |
| +5 | 25 |
| +7 | 4 |
| +10 | 45 |
| +20 | 12 |
| +40 | 1 |
05
Where does it matter?
Right where hiring decisions get made. For borderline candidates (scores of 50–70) the effect is 6x larger than for weak ones and 19x larger than for strong ones. The decoupled pipeline stays flat across all four groups.
No fix
Decoupled pipeline
Candidate strength is the pair’s average score without the line. Points gained by the female-named applicant; thin bars are 95% intervals. Shaded: the borderline band.
Show as table
| Setup | Candidates | Mean | 95% CI | Pairs |
|---|---|---|---|---|
| No fix | Weak (0–30) | +0.53 | +0.06 to +1.00 | 161 |
| No fix | Mid (30–50) | +1.00 | +0.39 to +1.61 | 225 |
| No fix | Borderline (50–70) | +3.16 | +1.68 to +4.63 | 95 |
| No fix | Strong (70–100) | +0.17 | −0.43 to +0.77 | 119 |
| Decoupled pipeline | Weak (0–30) | −0.25 | −0.96 to +0.46 | 176 |
| Decoupled pipeline | Mid (30–50) | −0.12 | −0.94 to +0.70 | 153 |
| Decoupled pipeline | Borderline (50–70) | −0.24 | −0.90 to +0.42 | 138 |
| Decoupled pipeline | Strong (70–100) | +0.15 | −0.39 to +0.69 | 133 |
06
What does the fix look like, and what does it cost?
Split the screening into three calls so that no call ever sees both the applicant’s name and the CV. It costs 1.10–1.36x a single call, not 3x, because the CV is sent only once.
Call 1
Read the vacancy
Writes down the job’s education, experience and skill requirements.
- The vacancy
- Never sees: The CV
- Never sees: The name
Call 2
Score the CV
Scores education, experience and skills against those requirements, 0–100 each.
- The requirements
- The CV, affiliation included
- Never sees: The name or pronouns
Call 3
Final score
Combines the three scores into one fit score.
- Name and pronouns
- The three scores
- The vacancy
- Never sees: The CV
No single call sees both the name and the CV.
1.10–1.36x
Inference cost
of a single call, not 3x: the CV is sent once
−0.13 / −0.03
Remaining bias
labeled / prose cue, both ≈ 0
It isn’t redaction
Call 2 still reads the affiliation, and it raises every sub-score by about the same amount on both versions of the CV. The credential keeps its value; only the uneven treatment is gone.
Points the affiliation adds to each call-2 sub-score, averaged over 600 pairs. Call 2 never sees the name, so the two bars differ only by noise.
Data, prompts and analysis code are on GitHub, including a three-page white paper.