2ndOpinion.FYI中文Log in
genius.wiki

#436 2015 · Avon and Somerset Police / Behavioural Insights Team · Public sector recruitment / assessment design

The test wasn't biased on paper, so the police force stopped looking at the test and looked at what happened right before it

the problem

A selection test showed a stark pass-rate gap between groups that content review of the test itself couldn't explain

background

Avon and Somerset Police's situational judgement test, part of its officer recruitment process, was producing a large pass-rate gap between Black and minority ethnic (BME) candidates and other candidates — roughly 30% versus roughly 65% by one account — even though the test's content had been reviewed and showed no overt bias in its questions or scoring. A test that looks fair on paper can still be shaped by psychological factors that have nothing to do with its content.

The Behavioural Insights Team, working with the force, hypothesized the gap was driven by stereotype threat — the well-documented effect where awareness of a negative stereotype about one's group can itself depress performance on a high-stakes evaluation, independent of actual ability. Standard fixes for test bias, like revising question wording or scoring rubrics, would do nothing to address a psychological effect happening before a candidate ever reads the first question.

what everyone would do

The standard response to a pass-rate gap on a selection test is to audit the test itself — reread the questions for biased language, check whether scoring weights disadvantage any group, and revise the content. That fails here because the gap wasn't coming from anything in the test; a content review had already found no overt bias, so revising questions or rubrics would have changed a document that wasn't the actual source of the disparity.

what they saw

The Behavioural Insights Team saw that a test can be perfectly fair on paper and still produce an unfair result, because performance depends on a candidate's psychological state walking in, not just the test's content. Stereotype threat — the documented effect where awareness of a negative stereotype about your group depresses performance on a high-stakes evaluation — operates entirely before the first question is read, which meant the fix had to happen upstream of the test, not inside it.

the move

Rather than change the test itself, the team added two brief additional questions to the email invitation candidates received before sitting the test, asking them to reflect on what would make them a good addition to the police force and what significance that would have in their community — a short self-affirming reflection designed to counter stereotype threat before the assessment began.

why it works

Adding two brief self-affirming reflection questions to the pre-test email — asking candidates what would make them a good addition to the force and what that would mean to their community — gives candidates a moment to reaffirm their own competence and belonging before the evaluation begins, which the stereotype-threat literature shows can buffer against the anxiety that otherwise depresses performance under stereotype-relevant pressure. Because the intervention targets the psychological state candidates bring to the test rather than the test's content or difficulty, it closes the pass-rate gap (a roughly 50% increase in pass probability for minority applicants, per the published field experiment, with no effect on other candidates) without lowering the bar or changing what's being assessed — the same test, taken by a candidate primed to feel they belong, produces a truer read of their actual ability.

the payoff

BME candidates who received the priming questions passed the situational judgement test at rates comparable to non-BME candidates, roughly closing a gap that had previously run from about 30% up to parity around 65%, without any change to the test's content, difficulty, or scoring.

where it breaks

The intervention only helps when the performance gap is actually driven by stereotype threat or a related psychological mechanism (belonging uncertainty, evaluation anxiety) rather than by a genuine skills or preparation gap, unequal access to practice materials, or bias embedded in the test's actual content — misdiagnosing the cause and applying a priming fix to a structural problem would produce no improvement while looking like an intervention was tried. It also depends on the priming being genuinely brief and low-stakes; an intervention that itself draws heavy attention to group identity or feels performative can backfire by making the stereotype more salient rather than less. And because the fix operates on psychology rather than content, it says nothing about whether the underlying test is actually measuring the right thing — an assessment could still have other validity problems the priming intervention would never surface.

what came after

The intervention is cited in UK public-sector behavioural science circles as a demonstration that fairness problems in assessment can originate in candidates' psychological framing rather than in test design, and the Behavioural Insights Team has since promoted similar priming-based approaches to other recruitment and assessment contexts facing unexplained demographic performance gaps.

references

  1. [1]Levelling the playing field in police recruitment: Evidence from a field experiment on test performancePublic Administration 95(4), 2017 (Linos, Reinhard & Ruda), via UC Berkeley Goldman School of Public Policy, 2017gspp.berkeley.edu
  2. [2]Better, fairer recruitment with behavioural scienceMoreThanNow, 2018morethannow.co.uk

keep it

same kind of clever

Back to the archive