2ndOpinion.FYIEN
genius.wiki

#1040 2007 · Galaxy Zoo (University of Oxford, University of Portsmouth) · Astronomy / citizen science

A million galaxy images would take years to sort by hand; astronomers gave it to strangers

问题

The Sloan Digital Sky Survey produced roughly a million galaxy images needing visual classification by human eye

背景

Automated image-processing software could locate galaxies in Sloan Digital Sky Survey data, but reliably classifying a galaxy's shape — spiral, elliptical, or merging — still required a trained human eye, and the survey had produced roughly a million such images. Assigning that classification task to graduate students, the standard approach, meant years of tedious manual work for a single researcher before the dataset could be used for any downstream science.

In July 2007, astronomers at Oxford and Portsmouth instead put the raw images on a public website, Galaxy Zoo, with a simple interface anyone could use to classify a galaxy's shape in seconds after a short tutorial, explicitly betting that enough public interest would substitute for specialist labor. They had projected that a few thousand engaged volunteers might get through the million images in a couple of years.

换别人会怎么做

The available option was assigning the classification task to graduate students or a small research team, which was the field's standard practice and would have taken years to work through a million images at any individual researcher's pace.

他们看到了什么

The team saw that classifying a galaxy's shape needs a trained eye, which a short tutorial can produce, not years of study, so the real bottleneck was labor-hours, not expertise, and the public had plenty to give.

那一手

Galaxy Zoo published its full set of roughly one million Sloan Digital Sky Survey galaxy images on a public website with a simple classification interface, letting anyone visually sort galaxies by shape after a brief tutorial, turning a specialist research task into a task any interested member of the public could contribute to directly.

为什么管用

Because the classification task could be taught in minutes and each image only needed a quick visual judgment, the project could recruit contributors with zero prior training and still trust the aggregate result, especially once multiple volunteers classified the same image and their answers were combined to cancel out individual error. The unbounded size of the potential volunteer pool, compared to a fixed research budget for graduate students, is what let the project finish in weeks a task originally estimated to take years.

值了多少

Volunteers logged 8 million classifications in the first ten days and over 50 million within a year, from 150,000 people.

什么时候会失灵

The mechanism depends on the classification task being genuinely learnable by a novice in a short tutorial and on being able to average out individual mistakes across many independent volunteer classifications of the same item; tasks requiring deep specialist judgment, or where a single classifier's error can't be caught by cross-checking against others, don't hold up as well under crowd classification.

后来呢

One volunteer, Dutch schoolteacher Hanny van Arkel, flagged an unusual object she couldn't classify that turned out to be a previously unknown type of astronomical phenomenon, later named Hanny's Voorwerp and published in a peer-reviewed paper with her credited; the project's success led to the broader Zooniverse citizen-science platform, which has since applied the same crowd-classification model to dozens of other scientific datasets.

资料来源

  1. [1]Hanny's Voorwerp: History of a mysteryZooniverse, 2013daily.zooniverse.org
  2. [2]Galaxy Zoo: 'Hanny's Voorwerp', a quasar light echo?Monthly Notices of the Royal Astronomical Society, 2009academic.oup.com

收下它