EN
Back to the archive

The encyclopedia · Engineering & Operations · Product decision · 1996–2001

The Internet Archive began saving the web before it knew how big the web would get.

In 1996 the Internet Archive started crawling and snapshotting web pages, so a future reader could still see what they looked like.

Internet Archive

the move

The Web changes constantly and pages disappear, so a link clicked today can be dead tomorrow, and the record of what a site once said is gone.

In October 1996 engineers at the Internet Archive launched crawlers that took snapshots of web pages, when the whole Web was about 2.5 terabytes.

Storing snapshots was cheap and the archive could not know in advance which ones would matter, so it recorded as much as it could and later exposed it through the Wayback Machine.

why it works

  • Saving a page costs far less than it costs to lose the only copy
  • Broad capture needs no judgment call about which pages are worth keeping
  • A time-indexed interface turns a pile of copies into a usable historical record
  • The archive preserves what institutions, companies and individuals on the Web once said
the payoffCapture everything now, decide value laterclever

what transfers

For irreversible loss, capture broadly rather than curate narrowly: keep every snapshot and let the future decide which one matters.

what came after

The Wayback Machine became an indispensable tool for historians, journalists, researchers and the public, and grew to capture hundreds of billions of pages by working with partners. The Internet Archive also faced legal challenges over copyright and preservation, while becoming a central target of debates about who should control access to knowledge.

references

spotted an error? The archive wants to know.

same kind of clever