2ndOpinion.FYIEN
genius.wiki

#877 2007 · Luis von Ahn's team, Carnegie Mellon University (Maurer, McMillen, Abraham, Blum) · Web security / mass digitization

He hid unreadable book words inside the CAPTCHA everyone was already solving

问题

Digitizing old books stalled wherever OCR failed — up to one word in five — and paying humans to retype was ruinous

背景

The mid-2000s mass-digitization boom — the Google Books Project, the nonprofit Internet Archive — photographed books from before the computer age and ran optical character recognition over the scans. On aged, degraded print, faded ink and yellowed paper, OCR failed on as many as one word in five in the Carnegie Mellon team's tests, and every failed word had to be retyped by a paid human. The backlog of unreadable words was the bottleneck of the whole enterprise, and it grew with every volume scanned.

Meanwhile an invention from the same university was minting exactly that labour for free. The CAPTCHA — the squiggly-word test devised at CMU in 2000 to keep automated programs out of web services — was being solved more than a hundred million times a day by people proving they were human, a few seconds each: hundreds of thousands of hours of human attention daily, spent reading characters a computer cannot read, and wasted by design.

换别人会怎么做

Every route treats transcription as labour to buy: pay human typists per word against a backlog of millions of volumes; pay engineers to improve OCR on print it has already proved powerless over; or digitize only clean modern volumes and leave the archive's oldest texts behind. Each prices the backlog at face value, and the backlog always wins.

他们看到了什么

A hundred million people a day proved they were human by reading characters computers couldn't — the exact skill the scanners lacked, burned as a security tax. The workforce was not missing — already on site, unpaid.

那一手

In 2007 the team launched reCAPTCHA. The test now serves two words: one lifted from a failing book scan, whose answer nobody knows, and one 'control' word whose answer is known. Type the control word correctly and you are certified human — and your reading of the unknown word is counted; once enough independent users agree on it, it is promoted into a control word in future puzzles. Verification and transcription became the same keystrokes, and within a year the system was running on more than 40,000 websites.

为什么管用

The control word made security and labour one keystroke: proving you were human by typing the known word was simultaneously the quality check on your transcription of the unknown one, so the added work cost almost no extra time and no wages. Agreement among strangers who never meet replaced paid proofreading — only answers that converged counted, which is how an anonymous crowd matched the 99-percent guarantee of professional transcription services. And every transcribed word became a new control word, so the system's security capital grew out of its own output: scaling the work hardened the gate instead of weakening it.

值了多少

Year one: 1.2 billion puzzles solved, 440 million words transcribed at better than 99% accuracy — 17,600 books' worth, at zero labour cost.

什么时候会失灵

The pattern needs three things at once and fails without any: a captive moment that repeats at scale for reasons of its own, micro-tasks with checkable right answers, and answers strangers converge on — legible words qualify; taste, judgement and context do not. The added burden must stay near zero, because the moment the embedded task noticeably slows the ritual it rides on, users abandon the ritual and both purposes die. And the labour is unpaid: defensible while the harvest is a public archive, on the strength of the effort being wasted anyway — a justification tied to the waste, not a licence for whatever the pattern is pointed at next.

后来呢

Google acquired reCAPTCHA in September 2009; by then it ran on more than 100,000 sites — Facebook, Craigslist and Twitter among them — and had helped digitize nearly all of the New York Times archive alongside the Internet Archive's books. It became the founding case of von Ahn's 'human computation' agenda, redirecting effort humans were already spending, which his lab extended to games that label photos and audio and which other researchers extended to tasks like protein folding with Foldit.

资料来源

  1. [1]reCAPTCHA: Human-Based Character Recognition via Web Security Measures, Science 321 (paper PDF)Science (AAAS), open copy via Semantic Scholar, 2008pdfs.semanticscholar.org
  2. [2]Computer users are digitizing books quickly and accurately with Carnegie Mellon methodEurekAlert (AAAS), Carnegie Mellon University release, 2008eurekalert.org
  3. [3]reCAPTCHA (a.k.a. Those Infernal Squiggly Words) Almost Done Digitizing New York Times ArchiveNewsweek, 2009newsweek.com

收下它

同一路聪明