The encyclopedia · Engineering & Operations · Product decision · 1996–2001
The Internet Archive began saving the web before it knew how big the web would get.
In 1996 the Internet Archive started crawling and snapshotting web pages, so a future reader could still see what they looked like.
Internet Archive
the move
The Web changes constantly and pages disappear, so a link clicked today can be dead tomorrow, and the record of what a site once said is gone.
In October 1996 engineers at the Internet Archive launched crawlers that took snapshots of web pages, when the whole Web was about 2.5 terabytes.
Storing snapshots was cheap and the archive could not know in advance which ones would matter, so it recorded as much as it could and later exposed it through the Wayback Machine.
why it works
- Saving a page costs far less than it costs to lose the only copy
- Broad capture needs no judgment call about which pages are worth keeping
- A time-indexed interface turns a pile of copies into a usable historical record
- The archive preserves what institutions, companies and individuals on the Web once said
what transfers
For irreversible loss, capture broadly rather than curate narrowly: keep every snapshot and let the future decide which one matters.
what came after
The Wayback Machine became an indispensable tool for historians, journalists, researchers and the public, and grew to capture hundreds of billions of pages by working with partners. The Internet Archive also faced legal challenges over copyright and preservation, while becoming a central target of debates about who should control access to knowledge.
references
spotted an error? The archive wants to know.