hckrnws
Japanese used bookstores see 5x sales surge as books are being bought by the ton
by speckx
by speckx
I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.
But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.
Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.
We were taught respect for books, but not everything is worth preserving.
Most of the information in it is practically worthless today unfortunately - Saddam Hussein still alive, as were Katharine Hepburn and King Hussein of Jordan, no mention of Ceres and Pluto was still a planet, and the Internet and software had barely a mention. Most articles had less information than a standard Wikipedia article.
That being said, the way information was presented in them still outshines anything one may find on the internet today. Even the simple elements - neat diagrams and relevant images, proper sectioning and organization of the text, footnotes to other relevant articles...
Long gone are the days when I would simply take a bowl of ice cream and a volume and just read it end to end.
I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are
Yeah, now instead of old books being turned into toilet paper, we'll get turned into toilet paper. Much better.
But Sam Altman will become richer than God, and isn't that what really matters?
But don't worry! You'll still have access to ChatGPT until your savings run out.
> It has been discovered that online used bookstores across Japan have been receiving a surge of large orders for books since around August of this year. Interviews with these bookstores reveal reports of "100 books sold per day" and "days where sales have increased fivefold," leading to widespread speculation within the industry that the orders are intended to collect training data for generative AI (artificial intelligence). Further investigation revealed records of over 50 tons of books being exported from Japan to the United States. Is it acceptable for books to be consumed and discarded for AI training?
[...]
> Nippon Television investigated using "Sayari," a tool that analyzes import and export data, and confirmed records that a group company of this major Japanese book distributor exported more than 50 tons of "JAPANESE BOOKS" to the US since last year. Assuming that all the books were heavy hardcovers (calculated at 500 grams), this would amount to the equivalent of 100,000 books.
[dead]
[deleted]
it seems really sad that these books data isnt open in the first place...
They didn't bother when it was a non-profit sponsored by Elon Musk.
1) The labs don't want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.
2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that's why the labs are starting to purchase by the pallet, it's because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.
3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don't get destroyed or thrown away.
Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don't expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.
[deleted]
[dead]
[dead]
[dead]
Training an LLM on copyrighted works is not illegal.
This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.
People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.
I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.
The relevant part of the ruling starts on page 27.
For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
The third fair use factor favors fair use for the purchased library copies converted from print to digital.
Fair use favors the destruction to avoid accidental surplus copying of the source material. It doesn't require it. If Anthropic put everything in a warehouse, they could keep it in the warehouse... but they couldn't do anything with them afterwards. They couldn't sell them as books or donate them as that would mean that it wasn't fair use and copyright infringement would have taken place once that additional copy was distributed again. At that point, fair use and economics both favor destroying the original.If I make a DVD copy of an old VHS tape that I own, that's fair use format shifting. I can keep the VHS tape without issue. I cannot donate it to the library or put it out in a garage sale. When I got rid of my VHS player, I threw out VHS tapes too since they were of no use and only cluttered my shelf.
I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.
I was responding to a comment, which apparently is another thing you missed. Maybe work on not doing that, instead of snarkily asking for everything to be carefully spelled out to you?
> or are you just mad it enriches people you don't like?
Nice strawman, pity if someone knocked it down.
Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.
The goal here is to have all human knowledge in a single file, which is pretty neat IMO.
I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.
Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.
They're scanning millions of books.
It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.
I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.
There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.
Lots and lots of information is not online. You'd be surprised.
Books are more likely to be about a specific topic or story or time or setting and be more information dense
LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.
They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).
[dead]
OCR Scanned for training, then tossed away or burnt. Great for nature.
Can we really trust these AI companies when they're basically assimilating human art and culture and everything, only to go on and destroy the evidence?
There's a general sense that librarians are preservationists. This is far from the truth. Most books get pulped within a few years of printing. Librarians are constantly discarding books that don't get checked out to make space for new books.
[deleted]
> the books being burned by these misanthropic lunatics are not available in any other medium
prove it, name one title
> this is not about fascination with some particular mediumn of transmission
it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.
if they were available, the labs would have purchased them digitally or pirated them
> Most books of value don't get destroyed or thrown away.
Of value to who? Most used bookstores are boutiques that over-curate and will happily refuse or recycle books that they deem are inferior/irrelevant. This is a big reason why I prefer Half Price Books over most any other used bookstore, they sell most everything.
Yeah, but like I said in the comment you are replying to, somewhere around 1-5% of books are worth the paper they are printed on and far more, and those are exactly the books the labs are trying to scan and destroy, they are not buying pallets of books and scanning them in order to scan the june 1995 tv guide for the 800th time, they want unique, interesting, and high quality text, because that's what feeds pretraining.
Lets find a random conference book on Amazon... Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference. It was published in 1959. It's a rare book in that I can only find one copy of it on Amazon.
https://www.amazon.com/Radiobiology-Radiology-Proceedings-Mi...
Lets go find it...
https://search.catalog.loc.gov/instances/e3fb7a94-2b25-5388-...
There's one copy onsite at the Library of Congress in the General Collections and another copy offsite. You can get a reading card for the Science and Business Reading Room ( https://www.loc.gov/research-centers/science-and-business/ ) and request that book and read it.
It also happens that my alma mater has a copy of the book too. https://search.library.wisc.edu/catalog/999554514302121 - it's in the stacks in Ebling Library.
Every book that has been published, there's a copy of it somewhere. If it was published in the UK, it is in the British Library ( https://youtu.be/ZNVuIU6UUiM ).
The books are there.
What's happening is that these books are getting bought by AI companies and scanned. If they weren't bought by the AI company then... the book in the... I'm not even sure. It's in Portland... if inventory doesn't sell, it may instead get thrown out. Maybe this year, maybe in ten years - it costs money to have inventory that doesn't sell.
Libraries will do book sales of books that aren't checked out frequently. https://www.booksalefinder.com . Sometimes they're donated, sometimes it's library discards.
What do libraries do as they cull their collections?
https://nwls.wislib.org/what-to-do-with-discarded-books/
> Some libraries will donate weeded materials to community resale shops, Goodwill, or other resale shops. Some will take all of your unwanted books without question, but many resale shops don’t have the capacity to accept the large number of books that libraries discard. If you do find a place to take them, you’ll still need to use staff or volunteer time to get them there.
> You’ll often hear this idea from well-meaning community members, and there are some fun crafts that can be made from discarded books. A quick online search will turn up many ideas for ways to use your books in crafts for kids, teens, and adults, including folded book art, blackout poetry, collage or prints made on book pages, wreaths and garlands made from book pages, and more. But even if you do a LOT of craft activities at your library, it’s unlikely that you could use up enough of your discards to even notice the difference.
For example... https://reallifeartist.wordpress.com/tag/book-carving/ or https://www.cutandfoldbookart.com/cut-and-fold-method-instru...
These are books that used book stores (and libraries) are trying to get rid of to free up room for inventory that will move or shelf space for books that people will check out and read.
... but the last copy of a book can always be found in the national library where it was published... though if you really want to preserve Genetics, Radiobiology and Radiology Proceedings, Mid-Western Conference from 1959 from being slurped up by LLM training, you can buy a copy and shelf it yourself.
scanning+ocr: 0.05 per page * 300 pgs: 15/bk
acquisition+disposal cost: 2.00/bk
total acq+scan+dispose = 425M
settlement value / bk: $3,000 [bartz v anthropic]
settlement probability, let's say 1%: 0.01
expected settlement / bk : $30
expected settlement total: 750M [industry-wide this is certainly an underestimate]
legal fees: 25% of settlement total: 187M
total 1.36B
+ operations, engineering, storage
> But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.
BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.
Both definitions of knowledge creation can used. If one creates a book and it is never read, has it been produced? Isn't knowledge also its access and how widely it is disseminated?
The benefit is that it properly mocks a misleading framing.
> If one creates a book and it is never read, has it been produced?
If one destructively scans a book that will never be read, has anything been lost?
Note: I made no personal attacks.
The book ceases to be a book right at the beginning of the extractive process.
It's not as though they particularly care, and it's not as though bad press matters with a captured regulatory environment.
https://www1.folha.uol.com.br/ilustrada/2016/09/1818064-se-n...
https://www.jornada.com.mx/2005/05/25/index.php?article=a04n...
https://www.lanacion.com.ar/cultura/a-donde-van-libros-no-ve...
https://revistadossier.udp.cl/dossier/destruir-con-permiso/
Looks to me, going by this reporting, that LATAM treats books identically to the US (The first headline, translated, is literally "If not sold or donated, books can become toilet paper")
Very little was gained.
> HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.
https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._HathiTr... and https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,...
> Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google's Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.
> The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.
The bill on orphan works that didn't pass was https://en.wikipedia.org/wiki/Shawn_Bentley_Orphan_Works_Act...
https://www.hathitrust.org/the-collection/search-access/copy...
And there are exceptions for copyrighted works allowing them to lend them out.
> Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...
Scanning still continues. https://www.hathitrust.org/member-libraries/contribute-conte... - though it's not at the same rate as it was during google books project.