#103 2002 · Chicago Public Schools (with researchers Brian Jacob and Steven Levitt) · Education / statistical fraud detectionlegibility
You can't catch a cheating teacher by watching one classroom — but you can catch them by watching the pattern across thousands
the problem
A single classroom's suspiciously good test results look exactly like a single classroom's honestly good test results — there's no way to tell fraud from genuine excellence by inspecting one case in isolation, no matter how closely you look
background
High-stakes standardized testing creates real incentive for teachers and administrators under pressure to inflate scores — changing answers, feeding answers during the test, or other manipulation — but any one classroom's unusually strong results are, viewed alone, indistinguishable from a genuinely excellent teacher having a strong year. Manual investigation of individual suspicious classrooms is slow, expensive, and easy to dispute, since a single score jump has an innocent explanation available almost every time.
Researchers Brian Jacob and Steven Levitt built a statistical algorithm that didn't look at any single classroom's result in isolation, but at two specific fingerprints only fabricated results tend to produce together: an unusually large year-over-year score jump for a given classroom, combined with suspicious block patterns of identical answers across students in that same classroom — a signature genuine student performance essentially never produces, since real students don't converge on identical wrong answers in the same unusual sequence.
the move
In spring 2002, Chicago Public Schools invited Jacob and Levitt to flag classrooms for its existing retest quality-control program based on the algorithm's output, splitting the 117 retested classrooms into three groups: those the algorithm flagged as likely cheating, 'good teacher' classrooms with large score gains but normal answer patterns, and a randomly selected control group — a design that let the algorithm's predictions be tested against real, monitored retest performance rather than remaining a statistical suspicion.
the payoff
On the closely monitored retest, algorithm-flagged classrooms saw scores collapse by more than a full grade equivalent, while 'good teacher' classrooms with genuinely strong initial gains held steady or even improved slightly, and the random control group showed only a small decline — a clean statistical separation between real performance and fabricated performance that manual classroom-by-classroom investigation could never have produced with comparable confidence or speed.
what came after
The Jacob-Levitt methodology is a foundational case study in applied statistics and fraud-detection literature, cited across education-policy and data-science curricula as a model for detecting manipulation through population-level statistical signatures rather than individual case review, and comparable answer-pattern and score-jump detection methods are now standard practice at testing organizations and state education departments.
references
- [1]NBER Digest — Do Incentives Cause Teachers to Cheat?National Bureau of Economic Research, 2003nber.org
- [2]The Cheating CurveChicago Booth Review, 2013chicagobooth.edu