#186 2002 · Chicago Public Schools (with researchers Brian Jacob and Steven Levitt) · Education / statistical fraud detection
You can't catch a cheating teacher by watching one classroom — but you can catch them by watching the pattern across thousands
the problem
A single classroom's suspiciously good test results look exactly like a single classroom's honestly good test results — there's no way to tell fraud from genuine excellence by inspecting one case in isolation, no matter how closely you look
background
High-stakes standardized testing creates real incentive for teachers and administrators under pressure to inflate scores — changing answers, feeding answers during the test, or other manipulation — but any one classroom's unusually strong results are, viewed alone, indistinguishable from a genuinely excellent teacher having a strong year. Manual investigation of individual suspicious classrooms is slow, expensive, and easy to dispute, since a single score jump has an innocent explanation available almost every time.
Researchers Brian Jacob and Steven Levitt built a statistical algorithm that didn't look at any single classroom's result in isolation, but at two specific fingerprints only fabricated results tend to produce together: an unusually large year-over-year score jump for a given classroom, combined with suspicious block patterns of identical answers across students in that same classroom — a signature genuine student performance essentially never produces, since real students don't converge on identical wrong answers in the same unusual sequence.
what everyone would do
Investigate individually suspicious classrooms one at a time -- pull test booklets, interview teachers, review score jumps case by case -- the standard fraud-response playbook, which is slow, expensive, and easy for an accused teacher to dispute, since any single classroom's good result always has an innocent explanation available.
what they saw
A score jump alone proves nothing, and neither does an unusual answer pattern alone -- but Jacob and Levitt saw that fabrication leaves a specific FINGERPRINT no single signal captures: an implausibly large gain paired with students converging on identical, often identically wrong, answer sequences. No genuine classroom produces both at once, so the combination, not either signal in isolation, is what actually distinguishes fraud from excellence.
the move
In spring 2002, Chicago Public Schools invited Jacob and Levitt to flag classrooms for its existing retest quality-control program based on the algorithm's output, splitting the 117 retested classrooms into three groups: those the algorithm flagged as likely cheating, 'good teacher' classrooms with large score gains but normal answer patterns, and a randomly selected control group — a design that let the algorithm's predictions be tested against real, monitored retest performance rather than remaining a statistical suspicion.
why it works
Genuine student improvement can produce a large score jump on its own, and coincidental answer-matching can happen in any classroom by chance, so either signal alone generates false positives that make case-by-case accusation unreliable and disputable. Requiring both signals together in the same classroom is a much higher bar that fabrication reliably clears (because someone changing or feeding answers naturally creates exactly this pattern) and honest performance almost never does by chance, which is why algorithm-flagged classrooms collapsed on a monitored retest while genuinely strong classrooms held their scores -- the algorithm wasn't guessing, it was reading a signature that only one underlying cause actually produces.
the payoff
On the closely monitored retest, algorithm-flagged classrooms saw scores collapse by more than a full grade equivalent, while 'good teacher' classrooms with genuinely strong initial gains held steady or even improved slightly, and the random control group showed only a small decline — a clean statistical separation between real performance and fabricated performance that manual classroom-by-classroom investigation could never have produced with comparable confidence or speed.
where it breaks
The method depends on the fraud actually leaving a detectable statistical fingerprint -- a more sophisticated cheater who varies which answers get changed, or who inflates scores by a smaller, less anomalous margin, can evade the specific signature the algorithm was built to catch. It also requires a genuine natural experiment or monitored retest to validate the flags against, since without confirming that flagged classrooms actually behave differently under controlled conditions, a statistical anomaly remains merely suspicious rather than proven, and flagging on statistical signature alone risks false accusations in the rare genuine classroom that happens to match the pattern by chance.
what came after
The Jacob-Levitt methodology is a foundational case study in applied statistics and fraud-detection literature, cited across education-policy and data-science curricula as a model for detecting manipulation through population-level statistical signatures rather than individual case review, and comparable answer-pattern and score-jump detection methods are now standard practice at testing organizations and state education departments.
references
- [1]NBER Digest — Do Incentives Cause Teachers to Cheat?National Bureau of Economic Research, 2003nber.org
- [2]The Cheating CurveChicago Booth Review, 2013chicagobooth.edu