Skip links

Amazon Allegedly Destroys Rare Texts for AI Training, Originating from Book Sales

Amazon has recently come under scrutiny for its acquisition of rare books, which it reportedly disassembles and digitally scans to enhance its artificial intelligence (AI) training initiatives. This revelation comes from an investigation by 404 Media, which tracked a rare book that ended up at an Amazon facility in Las Vegas using a hidden tracking device.

The facility in question, referred to as VGT3, features a unique symbol: a dinosaur grasping a book in its claws. In response to the investigation, Amazon stated that it sources books through standard commercial methods to improve the products and services available to its customers.

Tech giants like Amazon require vast amounts of textual data to train their large language models (LLMs). These models have already processed a significant amount of publicly available online content and, in some controversial cases, have utilized illegally obtained materials. Rare books—especially those no longer in print or hard to locate on the internet—provide an intriguing new source of high-quality training data.

The value of these texts lies in the fact that anything published prior to 2022 wouldn’t have been generated by an AI model, thus ensuring originality. Training on AI-generated text can lead to what is known as “model collapse,” a phenomenon where the quality of an LLM’s outputs deteriorates due to overexposure to inferior AI-generated content.

Editor’s Take

This development raises critical ethical questions about the treatment of rare literary works and their role in AI development. For users and businesses alike, the implications could be significant, potentially affecting the quality and authenticity of AI-generated content. As the AI industry continues to evolve, balancing innovation with respect for intellectual property will be crucial.

Source: techcrunch.com

Leave a comment