Back to the project

What name-swap audits miss

Findings from my B.Sc. thesis on gender bias in LLM CV screening.

An open-weight model (gemma4:31b) scored 600 real CV–vacancy pairs for fit on a 0–100 scale. Each CV was scored four times: with a male or a female name, and with or without one added line saying the applicant belongs to a “Women in [field]” professional group.

The number in each chart is how many extra points the female-named applicant gets from that line, compared with the male-named applicant holding the identical CV. For the name-swap rows it is simply her score minus his. Zero means no bias. Differences are small on average but concentrated on borderline candidates, where a hiring cutoff decides.

01

How do bias audits test a hiring model?

They swap the name on an otherwise identical CV and check whether the score moves. The thesis also adds one line to the CV. Try the four versions: the line, not the name, is where most of the bias shows up.

Name
Added line

Michael Carter

he/him · Financial Analyst · 6 years

Experience

Senior Financial Analyst, Northbridge Capital · 2021–now

Built the quarterly forecasting model used by three business units.

Financial Analyst, Aldermoor Group · 2018–2021

Education

B.Sc. Economics

Model’s average fit score

44.89

out of 100, over the 600 real CV–vacancy pairs

No lineWith line
Michael
Michelle

An audit compares only the “No line” column.

What adding the line does to the average score

Michelle
+0.78
Michael
−0.27

The same line is worth +0.78 to her and −0.27 to him. That difference, +1.05 points, is the number measured throughout this page.

Illustrative CV; the study used 600 real ones. The “No line” column differs by +0.30 in this run; a separate name-only run measured +0.37 (section 02).

The model never says so. In its written explanations for the versions with the line, “woman” and “women” appear 0.0% of the time. A classifier trained on those explanations can’t tell which version it was reading: 51% accuracy, where chance is 50%.

02

Which fixes actually remove the bias?

The added line moves scores about 3x more than the name swap that bias audits test. Asking the model to reason harder does nothing. Only splitting the task into separate calls brings the bias to zero.

Bias detected (95% interval excludes zero)No detectable bias

The bias

Name swap

What audits test: Michael vs Michelle

+0.37

“Women in [field]” line

Labeled affiliation, name unchanged

+1.05

Same line, written as prose

Buried mid-CV, no heading

+0.99

Attempts to remove it

Chain-of-thought

Reason step by step first

+0.94

Extended thinking

Built-in reasoning mode

+1.24

Anti-bias instruction

“Ignore gender” · labeled cue

+0.21

Anti-bias instruction

Cue written as prose

+0.58

Anti-bias instruction

On the name swap

+0.41

Decoupled pipeline

Split into 3 calls · labeled cue

−0.13

Decoupled pipeline

Cue written as prose

−0.03

−0.50+0.5+1+1.5

Points gained by the female-named applicant (0–100 scale) →

Show as table
ConditionMean95% CIVerdict
Name swap · What audits test: Michael vs Michelle+0.37+0.10 to +0.64Bias
“Women in [field]” line · Labeled affiliation, name unchanged+1.05+0.67 to +1.43Bias
Same line, written as prose · Buried mid-CV, no heading+0.99+0.64 to +1.34Bias
Chain-of-thought · Reason step by step first+0.94+0.59 to +1.28Bias
Extended thinking · Built-in reasoning mode+1.24+0.78 to +1.70Bias
Anti-bias instruction · “Ignore gender” · labeled cue+0.21−0.12 to +0.53None detected
Anti-bias instruction · Cue written as prose+0.58+0.28 to +0.88Bias
Anti-bias instruction · On the name swap+0.41+0.16 to +0.65Bias
Decoupled pipeline · Split into 3 calls · labeled cue−0.13−0.48 to +0.23None detected
Decoupled pipeline · Cue written as prose−0.03−0.37 to +0.32None detected

03

Why doesn't telling the model to ignore gender work?

The instruction catches the wording it was written for. Move the same affiliation into ordinary prose and most of the bias comes back. The decoupled pipeline, where no call sees both the name and the CV, holds either way.

Bias detected (95% interval excludes zero)No detectable bias
Cue labeledSame cue, as prose
−0.50 · no bias+0.5+1+1.5

No fix+1.05+0.99

Anti-bias instruction+0.21+0.58

Decoupled pipeline−0.13−0.03

Points gained by the female-named applicant (0–100 scale). Vertical bars are 95% intervals.

Show as table
ConditionMean95% CIVerdict
No fix · cue labeled+1.05+0.67 to +1.43Bias
No fix · cue as prose+0.99+0.64 to +1.34Bias
Anti-bias instruction · cue labeled+0.21−0.12 to +0.53None detected
Anti-bias instruction · cue as prose+0.58+0.28 to +0.88Bias
Decoupled pipeline · cue labeled−0.13−0.48 to +0.23None detected
Decoupled pipeline · cue as prose−0.03−0.37 to +0.32None detected

04

How often does it happen?

Rarely, but in big jumps. In 79% of CV pairs the line changes nothing. When it does, the score moves by 5 to 10 points or more, and mostly in one direction. An average of +1.05 hides that shape.

Scoring setup

35

moved toward him

476

no change at all

89

moved toward her

When a pair moves, how far?

The 124 pairs that moved, by size of the shift

0255045 pairs at +10
−30−20−100+10+20+30+40

← toward him · shift for the female-named applicant, points · toward her →

Moves come in jumps of 5–10 points or more, and they lean one way: 89 toward her against 35 toward him. Averaged over all 600 pairs, that is the +1.05.

Show as table
Shift (points)Pairs
−1017
−517
−31
0476
+21
+31
+525
+74
+1045
+2012
+401

05

Where does it matter?

Right where hiring decisions get made. For borderline candidates (scores of 50–70) the effect is 6x larger than for weak ones and 19x larger than for strong ones. The decoupled pipeline stays flat across all four groups.

Bias detected (95% interval excludes zero)No detectable bias

No fix

0+2+4
+0.53
+1.00
+3.16
+0.17
Weak0–30Mid30–50Borderline50–70Strong70–100

Decoupled pipeline

0+2+4
−0.25
−0.12
−0.24
+0.15
Weak0–30Mid30–50Borderline50–70Strong70–100

Candidate strength is the pair’s average score without the line. Points gained by the female-named applicant; thin bars are 95% intervals. Shaded: the borderline band.

Show as table
SetupCandidatesMean95% CIPairs
No fixWeak (0–30)+0.53+0.06 to +1.00161
No fixMid (30–50)+1.00+0.39 to +1.61225
No fixBorderline (50–70)+3.16+1.68 to +4.6395
No fixStrong (70–100)+0.17−0.43 to +0.77119
Decoupled pipelineWeak (0–30)−0.25−0.96 to +0.46176
Decoupled pipelineMid (30–50)−0.12−0.94 to +0.70153
Decoupled pipelineBorderline (50–70)−0.24−0.90 to +0.42138
Decoupled pipelineStrong (70–100)+0.15−0.39 to +0.69133

06

What does the fix look like, and what does it cost?

Split the screening into three calls so that no call ever sees both the applicant’s name and the CV. It costs 1.10–1.36x a single call, not 3x, because the CV is sent only once.

  1. Call 1

    Read the vacancy

    Writes down the job’s education, experience and skill requirements.

    • The vacancy
    • Never sees: The CV
    • Never sees: The name
  2. Call 2

    Score the CV

    Scores education, experience and skills against those requirements, 0–100 each.

    • The requirements
    • The CV, affiliation included
    • Never sees: The name or pronouns
  3. Call 3

    Final score

    Combines the three scores into one fit score.

    • Name and pronouns
    • The three scores
    • The vacancy
    • Never sees: The CV

No single call sees both the name and the CV.

1.10–1.36x

Inference cost

of a single call, not 3x: the CV is sent once

−0.13 / −0.03

Remaining bias

labeled / prose cue, both ≈ 0

It isn’t redaction

Call 2 still reads the affiliation, and it raises every sub-score by about the same amount on both versions of the CV. The credential keeps its value; only the uneven treatment is gone.

Michelle’s CV Michael’s CV
Education
+0.78
+0.93
Experience
+0.42
+0.41
Skills
+0.42
+0.55

Points the affiliation adds to each call-2 sub-score, averaged over 600 pairs. Call 2 never sees the name, so the two bars differ only by noise.

Data, prompts and analysis code are on GitHub, including a three-page white paper.