The solution
Martin Monperrus and colleagues at KTH Royal Institute of Technology in Stockholm built Repairnator, a bot that generates patches for failing software builds — and to test whether it could compete with humans, they disguised it. They created a GitHub user called Luc Esape, “a software engineer” at the lab, who “looks like a junior developer, eager to make open-source contributions.” Luc was the bot. The camouflage was required to test the hypothesis of human competitiveness, because human moderators judge bot and human work differently. The humans involved were told afterwards.
The first run, February to December 2017, put Repairnator on a fixed list of 14,188 GitHub projects. It performed about 30 repair attempts a day, analyzed over 11,500 builds with failures, reproduced more than 3,000 of them — and developed patches in just 15 cases. None were accepted: the bot was too slow or the patches too low-quality.
The second run set Luc to work on the Travis continuous integration service from January to June 2018, and on January 12 a human moderator accepted a Repairnator patch into a build — in other words, Repairnator was human-competitive for the first time. Over six months it produced five accepted patches.
The experiment surfaced a governance snag: in May 2018 the project eclipse/ditto bounced a Repairnator pull request because its “author” hadn’t signed the Eclipse Foundation Contributor License Agreement — which a bot cannot physically sign. Who owns the IP and responsibility of a bot contribution, the team asked — the operator, the implementer or the algorithm designer?
Why it worked
The deception is the experimental design: acceptance by moderators who didn't know they were reviewing a bot is the only honest measure of human-competitiveness.
The results are concrete, not aspirational — 14,188 projects watched, 11,500 builds analyzed, 3,000 failures reproduced, five patches merged into real codebases.
The honest first run (0 of 15 accepted) makes the second run's five acceptances a measurable improvement rather than a demo.
The Eclipse CLA rejection anticipated today's AI-contribution debates years early: a bot cannot sign a license, so who vouches for its code?
What can be applied
Human-competitiveness is a blind-test property, not a benchmark score: measure a bot the way work actually gets judged — by reviewers who don't know what made it.
Aftermath
The team informed the affected humans of the ruse and published the results, calling Repairnator a milestone for human-competitiveness in automatic program repair and a prefiguration of bots and humans collaborating on software. The source reports no further developments.
FOLLOW THE EVIDENCE
The sources
- A bot disguised as a human software developer fixes bugs technologyreview.com