The solution
In a 2012 New York Times Bits piece, Quentin Hardy credits part of Facebook's success to performance engineering. Friendster, founded in 2002, suffered frequent outages and slow page loads under heavy usage, and, in keeping with the practice of the time, put lots of information into a single database.
Facebook avoided this, in the article's words, by accident. It started at Harvard, then rolled out to other Ivy League schools, other colleges and then the public. Mark Zuckerberg created a separate database for each college and wrote software to connect people between them if they wanted to communicate. That model survives as many separate databases, now called shards, to which new users are assigned based on what the overall system can handle.
The article also credits a simplification, according to an executive who spoke anonymously: Facebook looked only at the mutual friends two people had, a much easier problem than Friendster's attempt to compute degrees of separation. Accel brought Jeff Rothschild, founder of Veritas Software, in as a consultant after investing in 2005, and Zuckerberg persuaded him to oversee the servers.
Why it worked
A single database for all users slows down and fails together as traffic grows.
Rolling out one school at a time let the team learn by doing and build anticipation.
Separate databases per college made failures local and the system more stable.
Counting mutual friends rather than degrees of separation kept the queries cheap.
What can be applied
A staged rollout can force an architecture that scales: partition by a natural unit before load makes you.
Aftermath
The article says Facebook's 900 million users sit in separate shards today. Friendster turned down a $30 million Google offer in pre-IPO stock in 2003, peaked at about 115 million members in 2008, then fell away and was sold in late 2009.
FOLLOW THE EVIDENCE
The sources
- Engineering Tricks That Helped Facebook Win archive.nytimes.com