SAN FRANCISCO — The Internet Archive runs on a paradox: it preserves the entire web for free while operating on donations and grants, storing over one trillion archived web pages in a model that works only if you never need to turn a profit.

Founded in 1996 by Brewster Kahle, the San Francisco non-profit digitizes and hosts websites, books, audio, software and video through archive.org. Its Wayback Machine—the searchable index launched in 2001—contains snapshots of nearly every public website ever crawled. The Archive also holds digitized books spanning 2003 to 2013, making them freely available.

The operation's scale is massive. Web crawlers, not user uploads, collect the majority of content. The Archive maintains multiple data centers and added BitTorrent distribution in 2012 to reduce bandwidth costs. International duplication followed: Internet Archive Europe opened in the Netherlands in 2004, and a Canadian mirror launched in 2016.

Kahle began the project in May 1996 while running Alexa Internet, a for-profit web crawling company he later sold to Amazon. The first archived page—a USA Today edition from December 1996—seeded what would become the largest digital library ever built by a non-profit.

The Archive generates revenue through specialized services. Archive-It contracts with institutions to crawl and preserve their sites. Open Library, its wiki-editable book catalog, attracts librarians and researchers. The Digital Accessible Information System (DAISY) format serves print-disabled users. NASA Images, housed on Archive servers, represents institutional partnership revenue.

Since 2017, the Archive has integrated its digitized book records into WorldCat, the global library catalog operated by OCLC, expanding discoverability. A digital arts residency program, launched in 2018, connects preservation work to cultural institutions.

The capital question looms: sustaining infrastructure at this scale—storage, bandwidth, redundancy, staffing—requires continuous funding. The Archive relies on donations, grants and service revenue. No advertising. No user data extraction. The model assumes that free access to knowledge is a public good worth subsidizing indefinitely, a bet that separates the Internet Archive from every commercial platform built in the past decade.