2ndOpinion.FYI中文
genius.wiki

#1040 2007 · Galaxy Zoo (University of Oxford, University of Portsmouth) · Astronomy / citizen science

A million galaxy images would take years to sort by hand; astronomers gave it to strangers

the problem

The Sloan Digital Sky Survey produced roughly a million galaxy images needing visual classification by human eye

background

Automated image-processing software could locate galaxies in Sloan Digital Sky Survey data, but reliably classifying a galaxy's shape — spiral, elliptical, or merging — still required a trained human eye, and the survey had produced roughly a million such images. Assigning that classification task to graduate students, the standard approach, meant years of tedious manual work for a single researcher before the dataset could be used for any downstream science.

In July 2007, astronomers at Oxford and Portsmouth instead put the raw images on a public website, Galaxy Zoo, with a simple interface anyone could use to classify a galaxy's shape in seconds after a short tutorial, explicitly betting that enough public interest would substitute for specialist labor. They had projected that a few thousand engaged volunteers might get through the million images in a couple of years.

what everyone would do

The available option was assigning the classification task to graduate students or a small research team, which was the field's standard practice and would have taken years to work through a million images at any individual researcher's pace.

what they saw

The team saw that classifying a galaxy's shape needs a trained eye, which a short tutorial can produce, not years of study, so the real bottleneck was labor-hours, not expertise, and the public had plenty to give.

the move

Galaxy Zoo published its full set of roughly one million Sloan Digital Sky Survey galaxy images on a public website with a simple classification interface, letting anyone visually sort galaxies by shape after a brief tutorial, turning a specialist research task into a task any interested member of the public could contribute to directly.

why it works

Because the classification task could be taught in minutes and each image only needed a quick visual judgment, the project could recruit contributors with zero prior training and still trust the aggregate result, especially once multiple volunteers classified the same image and their answers were combined to cancel out individual error. The unbounded size of the potential volunteer pool, compared to a fixed research budget for graduate students, is what let the project finish in weeks a task originally estimated to take years.

the payoff

Volunteers logged 8 million classifications in the first ten days and over 50 million within a year, from 150,000 people.

where it breaks

The mechanism depends on the classification task being genuinely learnable by a novice in a short tutorial and on being able to average out individual mistakes across many independent volunteer classifications of the same item; tasks requiring deep specialist judgment, or where a single classifier's error can't be caught by cross-checking against others, don't hold up as well under crowd classification.

what came after

One volunteer, Dutch schoolteacher Hanny van Arkel, flagged an unusual object she couldn't classify that turned out to be a previously unknown type of astronomical phenomenon, later named Hanny's Voorwerp and published in a peer-reviewed paper with her credited; the project's success led to the broader Zooniverse citizen-science platform, which has since applied the same crowd-classification model to dozens of other scientific datasets.

references

  1. [1]Hanny's Voorwerp: History of a mysteryZooniverse, 2013daily.zooniverse.org
  2. [2]Galaxy Zoo: 'Hanny's Voorwerp', a quasar light echo?Monthly Notices of the Royal Astronomical Society, 2009academic.oup.com

keep it