
This week, Publisher’s Weekly reported that Ingram Content Group sent an email to the publishers using their services saying that AI companies had begun buying physical copies of published books for the purposes of scanning and scraping the content. Why? To make use of the “fair use” caveat built into most copyright systems.
But to be clear, while fair use may apply to the person who buys a physical CD, rips the music to their computer, and then creates a playlist from that music to share with a handful of friends, it does not apply to commercial use. Meaning, if that person copied the CD they bought and then put those copies up for sale.
AI companies would argue that it’s perfectly legal for them to purchase a book, scrape the written content from it, and then feed it to their AI models. This may be true, since they’re using the content for “research” purposes. What’s done with it after the research is where they get into a legal gray area. If they’re allowing their LLM models to regurgitate Jurassic Park to anyone who asks their chatbot, then it’s a potential intellectual property lawsuit. And I would encourage anyone who finds that a chatbot is reciting passages from their book to internet users to immediately contact an IP lawyer.
Until AI companies learn to stop scraping people’s intellectual property without their permission, Ingram has created an opt-out form to combat this practice of purchasing physical books for AI “research”. Publishers (and indie authors) can add their name to the “do not sell” list. Ingram will take steps to try to keep AI companies from purchasing physical books directly from them. In some cases, the company might not know who is buying the books (for example, if an AI company has an employee order a book from Amazon), and therefore cannot interrupt the sale.
For more on this story, read the Publisher’s Weekly article. If you want to go directly to the opt-out form, click here.