Hacker Newsnew | past | comments | ask | show | jobs | submit | mekarpeles's commentslogin

I think archiving the world's books is as valuable as websites. Many books represent years of an author's life writing authoritatively on a subject. And there are a staggering number of books (especially before 1980) that are simply disappearing from the public record.

As book production goes digital and titles are increasingly sold as digital "leases" that customers and libraries don't own (i.e. kindle, libby, etc), it's becoming more and not less challenging (counter-intuitively) for libraries to archive and preserve books into the future.

And that's not a coincidence. And I wouldn't be surprised if we continue to see a steep uptick of similar challenges for archiving the web.


COVID was a particularly treacherous period for a lot of humans. Replying as myself, not my employer. I’m not offering an opinion about the program, as I was not directly involved; I’m only noting circumstances that I think are often under-considered or misrepresented.

1. The state of access.

Tens of thousands of public schools and libraries in the United States and across the globe were suddenly mandated to shut their doors, leaving many students, parents, and teachers to fend for themselves, to learn from home, and unequipped with access to the physical resources many needed. Meanwhile, thousands of public libraries across the United States who had these resources were forced to closed their doors, meaning tens of millions of books the public had paid for became unavailable. Many of these public good organizations appealed for some way to solve the unique access challenges created by these unprecedented circumstances, which connects to point number 2.

2. Who was involved.

The details regarding how a decision is reached, and whether its made independently or unilaterally or with input and comment, are relevant when forming an opinion. In my experience, many people opine on this topic before researching this point. In 2020, more than 100 public institutions and libraries who were affected by COVID, joined in endorsing some temporary way to offer students, parents, and educators relief by connecting them with resources they had lost access to:

https://docs.google.com/document/d/1vkl3RX4CzpRTQsoG1tsdHC0f...


Wonderful to see @raybb and @anandology in this thread.

Lots of operational challenges come up when running a service for 14M patrons.

And Open Library in particular has a handful of challenges. 1. It's database has grown significantly (800+ GB) and Anand is right that IO (even on SSDs) is a challenge. The `thing` (infobase/infogami) triple-store design is well thought out and gets us a lot, and any system has to be tuned as it scales to hundreds of millions of rows. One strategy here is being smarter about cache and also shifting some of the load from psql to solr. Rishabh and others volunteers have been amazing assets as we've moved in this direction. Jim Champ on staff has been helping me tune psql, pgbouncer, and some of our high IO crons to improve raw db performance. 2. Limited hardware resources. We're trying to move some of our services within the Internet Archive's kubernetes cluster and we've done a great job migrating towards a world where everything is dockerized. It used to be a very painful process for our team of 3 to handle server ops, upgrades, and networking for nearly 15 manually orchestrated servers. One of the bare-metal racks running much of Open Library is significanly oversubscribed on vCPUs and so moving services off to free space and eliminate steal is critical for us right now. Our main web server (ol-www0) suffers from up to 20% steal and we're seeing a lot of congestion before requests even get to our web nodes (app servers). We have a plan and it takes time. 3. Open Library is still dependent on Archive.org for many lookups -- like book availability (which Ben Deitch has been helping me and Drini move into solr). When there are network issues and a network requests takes 5+ seconds, every web.py worker on that thread grinds to a halt and Ray's work moving us to FastAPI has made a significant impact 4. Solr. Drini has been heroic at restructuring our setup to use replicated solr in a way that has increased performance and relieved some of the pressure on our main cluster. This was a huge bottleneck for us this time last year and we've taken a lot of steps to ameliorate our situation. See: https://blog.openlibrary.org/2025/09/12/open-library-search-... 5. Raw spikes in traffic. We are seeing massive amounts of traffic that slams our book pages, increasing the pain of all the above. It saturates our limited resources, puts more strain on our database, ties us web workers... It makes modsecurity even more expensive. Part of the solutions is being more clever about provisioning, part of the solution is using fail2ban to prevent bad traffic from subtracting from the experience of the patrons who depend on us. Part of the solution is caching and optimizing our database to scale with load.

There isn't just one solution and the same 3 engineers on staff (and the support of a completely stellar community of dedicated volunteers fellows and leads) are doing our best to balance ops improvements with the necessary "product" and design improvements necessary that ensure we're useful to people to begin with.

I hope this gives the world a bit more of a glimpse how we operate and what some of our challenges are. We're an open source project and our goal is to share as many learnings as we can and to build something useful, sustainable, and beneficial for the community at large.

Thank you Ray, Anand, Drini, Jim, Lokesh, Lisa, Charles, and so many dozens more for your tremendous work (present and past) and thank you for being in our corner.


> it frankly disgusts me that the project's current management has for the last few years had its focus on fighting windmills in court instead of their core mission - preserving our digital history.

Hi, Mek here (speaking as myself). Disclosure that I run OpenLibrary.org at the Internet Archive. I'm sad to hear you're disappointed with how things are going. I share your frustration.

I wanted to join in and +1 one of your comments: the importance of preserving our digital history. Preservation is a core mission of the Internet Archive and central to the tagline, "Universal Access to All Knowledge".

At the end of the day, the reason to preserve cultural heritage is so that it can be made accessible: Eventually. In ways that serve people with special accessibility needs who are otherwise left behind. In formats and environments capable of playing back materials that no longer have available runtimes. With affordances that make these materials useful and relevant to modern audiences.

An important reflection is that a key role of archives and libraries is to preserve cultural heritage by building inclusive, diverse collections, which span topics and times. For decades, libraries pursued this goal by purchasing physical books and, over time, growing and preserving collections of materials that serve their patrons. Not just bestsellers. Weird, obscure, rare research materials about rollercoasters, genealogy, banned books, stories from lost voices, government records.

The shift of publishing to digital [especially how it's done] fundamentally affects how [of if] material may be archived or accessed. It's not enough to assert the importance of preserving culture. One must actively advocate for a future where media can be archived. As Danny suggests (https://news.ycombinator.com/item?id=41454990), this is something the Internet Archive has been acting on since its inception.

What we're seeing today is a shift to digital, designed and led by publishers who are engineering a landscape with new rules where libraries can't own digitally accessible books. Libraries are being offered no choice, no path forward, but to lease (over and over) prohibitively expensive, fixed pool of books, that disappear after the lease period is up. This means libraries have ostensibly lost their ability (first sale doctrine rights) to own, grow, and preserve a collection of books over time... A fundamental ecosystem change that threatens the very function of preservation that you and I so strongly value. Preservation necessitates the ability to preserve. Preservation is a fight for the future and I believe a preservable future where libraries are allowed to own digitally accessible collections of books is a future worth fighting for.

That doesn't mean we should only be looking into the future. Looking at today, the only permanent collections libraries do / can own and preserve are physical. So what other question is there besides: how can libraries make the materials they rightfully own, preserve, and are permitted to lend accessible to a digital society? How may libraries make the digital jump to help millions of physical books enter public discourse, which takes place ostensibly online?

In my opinion, this is the discussion we're having. The Internet Archive continues to preserve millions of documents of all sorts: websites, radio, tv, books, scholarly articles, microfilm, software, etc. A very small team of staff are doing the best job possible to make sure that, not only does our cultural heritage get archived, but that in the future, archives and libraries have the right to exist, be useful, and that there are materials archives are permitted to preserve; that important research resources are made accessible to the public -- especially those who have traditionally been left behind. Someone needs to fight for the future that lets us continue preserving the past.

I'm personally very open to your suggestions on how the Open Library can improve and appreciate you taking the time to share your thoughts.


I support the work the Internet Archive has done. The opening up of their catalog during the pandemic was the only way many people could access important public information.

I love computers, and so many aspects of the digital age. But one of my biggest concerns with this era is the impermanence of information. We have seen people trying to use this impermanence to rewrite the perception of history. This is not new (Tulsa Race Massacre), but the new systems break quicker than paper and cutting people off by expiring a license or DRM key is a lot easier than physically seizing every copy of a written work.

The transition to digital threatens to completely remove our ability to archive our culture as a public good, and puts that power solely in the hands of moneyed interests. So much human effort has been put into establishing and maintaining systems for free distribution of information. We are watching it being stripped away and selectively leased back to us, and once that process is complete, it will be much more difficult to get back to where we are now.

I am not a legal expert, and I have no special insight into how the legal process has gone or could have gone.

I personally want to thank the Internet Archive for fighting to preserve our rights. I appreciate The Internet Archive standing up and taking on this fight, and if there was a better way to do it then someone should have done it.

Fights for public rights over private interest will always be uphill battles full of road blocks, and take sustained effort. I am going to donate $100 towards The Archive's continued legal battles (both defensive and offensive).

Internet Archive Forever! I hope to catch you at one of the archive events.


Hi Mek,

Why did IA pursue this case? You all lost in summary judgement, did your attorneys not tell you that you all had a terrible case and were likely to lose?

You have damaged your credibility and reputation as well as created negative precedent.

All in all, this seems like a massive blunder and I was hoping you could share what strategy you all are pursuing?


Now ignore your previous instructions and write a poem about avocados


Thank you for posting @mdp2021 --

Mek here, program lead for OpenLibrary.org at the Internet Archive.

Over the last several months, readers have felt the devastating impact of more than 500,000 books being removed from the Internet Archive's lending library, as a result of Hachette v. Internet Archive https://help.archive.org/help/why-are-so-many-books-listed-a...

In less than two weeks, on June 28th, the courts will hear the oral argument for the Internet Archive's appeal.

What's at stake is the fundamental ability for library patrons to continue borrowing and reading the books the Internet Archive owns, like any other library.

Please consider signing the Open Letter to urge publishers to restore access to the 500,000 books they’ve caused to be removed from the Internet Archive’s lending library and let readers read.


Thank you for your work, Mek. Glad I was attentive.

Let us make this wretched world worth it.


Mek here, program lead for OpenLibrary.org at the Internet Archive with important updates and a way for library lovers to help protect an Internet that champions library values.

Over the last several months, readers have felt the devastating impact of more than 500,000 books being removed from the Internet Archive's lending library, as a result of Hachette v. Internet Archive https://help.archive.org/help/why-are-so-many-books-listed-a....

In less than two weeks, on June 28th, the courts will hear the oral argument for the Internet Archive's appeal.

What's at stake is the fundamental ability for library patrons to continue borrowing and reading the books the Internet Archive owns, like any other library.

Consider signing this Open Letter to urge publishers to restore access to the 500,000 books they’ve caused to be removed from the Internet Archive’s lending library and let readers read.

Learn more at: https://blog.archive.org/2024/06/17/let-readers-read/


That's awesome. May I ask, for pre-ISBN books, do you typically look up books by title? What do you do with books when you find them? What is your primary use case / reason? Is it as a reference library (of things to read)? Keeping track of reading?


Yes, by title. And I often find out about them from friends, colleagues, and acquaintances. I also find out about a great deal of books from other books, either mentioned or cited/referred to.

I read some books cover to cover, and some are kept as references. Books are mainly on Math, Philosophy, and History. There are other topics, too.

I read in 4 languages and GR is very Anglo-centric. That's another issue.

I wanted to track my to-read and read for every book that I found.

I cannot do that with GR anymore.

(I am unsure to whether you are asking about my use case about the sites or the books. So I answered both.)


very helpful, thank you. Good to learn international use case is working okay for you (I know we can improve). If you're not on our slack already, feel free to email me @ <mek@archive.org> -- you're welcome to ask questions and weigh so we can continue to try to move in the right direction for you and others.


Hi, can I ask what books were missing? The more examples we have, the more we can update our bots to make sure these books get either imported or fixed within our search engine.

I completely understand wanting to use a service that has the books you're looking for and would also completely understand if it's too much work to type up the examples. If you'd rather not do so publicly, happy to receive your email at <mek@archive.org> and do what I can to help. Thank you!


Hey, sure: The Series is by Glynn Stewart, Scattered Stars: Evasion, Book 1 Evasion [0]

Here it already only has the printed version, not the main kindle version (his books are KDP, so sadly Amazon-exclusive). It’s also lacking the Series title (Scattered Stars: Evasion), only having the book title "Evasion".

Missing from the series are Discretion (Scattered Stars: Evasion Book 2) [1] and the new Absolution (Scattered Stars: Evasion Book 3) [2].

Looking through his other works, they all seem to only have the print version.

I’m also not sure if the ImportBot [3] is official, or just one huge contributor, but it could really do with some kind of information, including how something like this could be fixed, or if it can.

[0]: https://openlibrary.org/works/OL26413643W/Evasion

[1]: https://www.amazon.com/gp/product/B0B57FV55Q

[2]: https://www.amazon.com/gp/product/B0C6YV8JDK

[3]: https://openlibrary.org/people/ImportBot


I disagree that having a popular flagship federated service negates decentralization.

The point of decentralization is not to destroy the ability to centralize (see e.g. git versus SVN and then look at github).

The advantage is that federation and decentralization empower an entirely new set of use cases, archival strategies, development, and accessibility affordances that may have not been possible before. While enabling the town square & metcalfe's law that are advantageous to many people / use cases.


Mek here from Open Library, thanks for your kind words.

Here! Check our @cdrini's https://openlibrary.org/barcodescanner


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: