AI firms are quietly buying and destroying millions of printed books to train their models

As an antiquarian book collector of titles for specific regions of the world and specific topics of interest I can offer the following:
1) Once the written book has been destroyed you have no proof AI is lying to you. This is the entire point of the movement by the AI machines. Countries are generating AI slop as news to be sourced by AI as fact. China and Israel have spent $millions to do this.
2) Many of the books on archive.org have been scanned by another donor and provided to the organization. Look at the data underneath the book description. The downside to this is many of the books that were once availalable on the site are now gone-for whatever reason. Rare books are not destroyed in the process of scanning.
3) Folks that depend on google, wikipedia, etc., for their sources of fact and data have been lied to for a very long time. This is obvious when trying to have a normal conversation or debate on topics of interest, obscure, unpopular views, etc.
AI is another tool to be used for good and bad, depending on the developer and the user.
 
It's as if you read a tiny part of the article and then rushed to comment. This isn't about Ebooks. It's about the avoiding training AI on AI generated books. The article is saying there wasn't any AI generated books before 2022.
99.99% of the books are not AI-generated, so you can avoid training AI on AI generated books by simply doing nothing.
On top of that, the article is about printed books, I don't think even one book that's AI-generated and printed at the same time exists.
 
Actually, this is likely SAVING books for the future, just in electronic form. The vast majority of these books would probably never be read and eventually discarded anyway.
 
No, it doesn't require a book to be destroyed in order to scan it.

It's a cost cutting measure, lazy and negligent.

The type of scanning that the Internet Archive does isn't scalable at the levels the AI companies are vacuuming up the information. It's less cost cutting, lazy, or negligent, it's doing the task fast and efficiently. Since they aren't destroying books that are the last known actual copy of those books, the loss of one copy is immaterial to the existence of the book in general. The AI firms have no reason to rescan the same book multiple times, so it is literally one copy of a given book that likely has thousands of other copies extant elsewhere. They aren't destroying information permanently by doing so.
 
It's as if you read a tiny part of the article and then rushed to comment. This isn't about Ebooks. It's about the avoiding training AI on AI generated books. The article is saying there wasn't any AI generated books before 2022.
If you read a tiny part of the article "The focus is largely on books published before 2022, before synthetic text became widespread online." before rushing to condemn my comment you would have understood it.
 
Does training an ai on a purchased book and then make profits on the Ai violate the rights of ownership of a book?
 
Publishers wouldn't sell them digital copies, so they bought physical copies and chopped the spines off to feed them into their industrial scanner, then sent the scrap off for recycling.

It's a predictable outcome really.
 
Calm down.

They are destroying a single copy of a physical book not wiping it from existence.

(Note the destruction is only for legal proof that they’re not reusing after training or reselling it.)

The problem is they're destroying rare only 1 copy left in existence as well, and that's serious. They could easily manipulate the information in the books if all physical copies no longer exist and change history and people's perception of this world (which has already been heavily manipulated and hidden)
 
Back