AI Firms Buy Old Books, Hide Watermarks in Text
AI companies are buying pre-2022 books to avoid AI slop and quietly watermarking generated text. Here's what it means for your content.
AI Summary
AI companies are buying pre-2022 books to avoid AI slop and quietly watermarking generated text. Here's what it means for your content.
Here’s a strange but true snapshot of where AI is right now. On one hand, major AI companies are spending real money buying up shipping containers full of old, printed books, specifically ones published before 2022, because those books are guaranteed to be free of AI-generated writing. On the other hand, those same companies are quietly baking invisible, undetectable signatures into the text their own AI models produce, so it can be traced back later.
Two different problems, same underlying cause: AI-generated content has gotten so widespread across the internet that AI companies themselves no longer fully trust the web as a source of clean information. And they’re taking drastic, sometimes uncomfortable steps to deal with it, both by hunting down pre-AI human writing and by tagging their own output so it can be identified after the fact.
If you’re a business owner, startup founder, or local business publishing content online, both of these stories matter more than they might seem at first glance. One tells you something important about where AI models are getting smarter (and where they might start plateauing). The other tells you something you need to know before you publish another word of AI-assisted content on your website.
Let’s break down both trends properly, what’s actually happening, why it’s happening, and what it means for your content strategy going forward.
Part One: Why AI Companies Are Buying Old Books By The Truckload
The Core Problem: Model Collapse
To understand why anyone would want a warehouse full of old paperbacks, you need to understand a real, documented phenomenon called model collapse. This happens when an AI model gets trained repeatedly on text that was itself generated by earlier AI models, rather than genuine human writing. Each generation of training data gets a little more synthetic, a little more repetitive, and a little further from how actual humans think and write. Over time, quality quietly degrades.
The problem is that the open web, which used to be a goldmine of human-written text for training AI models, is now flooded with AI-generated content itself. A fast-growing share of new text published online today is machine-made, some of it low-effort “AI slop,” some of it more polished but still synthetic. For AI companies trying to train their next generation of models, that’s a real supply problem: where do you find large volumes of guaranteed human-written text anymore?
The Answer: Books Published Before 2022
This is where a company called ISBNdb enters the picture. ISBNdb, which describes itself as running the world’s largest book database, has been openly brokering bulk purchases of printed books for AI companies, specifically pitching books published before 2022 as the last large, reliable pool of text that’s structurally guaranteed to be free of AI contamination.
The logic is straightforward. Large language models didn’t exist at meaningful scale before 2022, so a book physically printed and published earlier than that simply cannot contain AI-generated text. As ISBNdb has put it in its own marketing, in a line widely quoted across tech coverage, books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.
How Large Is This, Actually
According to reporting, these deals aren’t small. AI firms have reportedly purchased anywhere from 1,000 to as many as 1 million pre-2022 books in a single engagement. That’s not a research library, that’s an industrial-scale data acquisition operation.
There’s a Second Reason Beyond Model Collapse: Data Poisoning
Beyond avoiding AI slop, there’s a second, more adversarial reason AI companies want clean, pre-2022 print sources. Some authors and artists have started fighting back against unauthorized AI training by deliberately poisoning their own text and images, using tools that embed patterns a human reader doesn’t notice but that quietly corrupt how a model learns from that content. Research cited by ISBNdb itself, drawing on Anthropic’s own findings, suggests that as few as 250 to 500 carefully crafted poisoned documents can plant a backdoor into a training corpus containing trillions of tokens. Pre-2022 print books were written before any of these poisoning tools existed, so they sidestep that risk entirely.
Quick Reference Table
| Aspect | Detail |
|---|---|
| Broker company | ISBNdb |
| Books targeted | Printed books published before 2022 |
| Volume per deal | Reportedly 1,000 to 1,000,000 books at a time |
| Main reason 1 | Avoiding AI slop and model collapse from synthetic training data |
| Main reason 2 | Avoiding intentional data poisoning tools like Nightshade |
| Confidentiality | Strict NDAs, buyer identities never disclosed |
| Legal precedent | Judge William Alsup ruled Anthropic’s book scanning and destruction practice was “clearly transformative” fair use |
| The controversial part | Scanning at scale typically destroys the physical book |
The Uncomfortable Part: These Books Get Destroyed
This is the detail that’s generated the most pushback, and it deserves a clear explanation.
Scanning a printed book at industrial scale, fast enough to process a million volumes, typically means slicing off the spine so loose pages can feed through a high-speed scanner. It’s faster and cheaper than any careful, non-destructive digitization method, but it physically destroys the book in the process.
ISBNdb has been candid about the optics problem this creates. Reporting has quoted the company’s own site acknowledging plainly that “AI company destroys two million books” isn’t a headline that generates sympathy. As a result, ISBNdb offers its AI company clients strict confidentiality: NDAs on every engagement, with buyer identities never disclosed publicly.
There’s also relevant legal context here. A federal judge, William Alsup, has already ruled that Anthropic’s specific practice of buying, scanning, and then destroying print books was “clearly transformative” fair use under copyright law. That ruling gives AI companies a meaningful legal foundation for continuing this practice, even as the ethics of mass book destruction remain an uncomfortable conversation.
Ingram, the largest book distributor in the United States, has reportedly warned publishers about this practice and is offering opt-out mechanisms. But the default remains that books can be scanned without explicit publisher permission, and the identities of the AI companies actually doing the buying are not disclosed.
Part Two: Why AI Companies Are Secretly Watermarking Generated Text
Now let’s flip to the other half of this story, because it’s directly connected. If clean, verifiably human-written text is becoming a scarce and valuable resource, it makes sense that AI companies would also want a reliable way to tag their own output, so that AI-generated text doesn’t quietly get mistaken for human writing and contaminate future training data, and so platforms and regulators can identify AI content when it matters.
What’s Actually Happening
Multiple major AI companies have quietly begun embedding invisible, machine-readable watermarks directly into the text their models generate. This isn’t something announced with a big product launch moment. It’s been rolling out gradually, largely disclosed through support documentation and legal pages rather than headline product announcements, with Gizmodo and other outlets piecing the rollout together after the fact.
Anthropic confirmed that new Claude models launched from August 2, 2026 onward embed this kind of watermark into every piece of generated text by default, applying worldwide. Google has been doing something similar with its SynthID technology for a while already, applied across text, images, audio, and video. OpenAI has used similar transparency tools for images and audio, though it hadn’t publicly detailed a text-specific detection system as of this writing.
How the Watermark Actually Works
The mechanism here is clever, and it explains why the watermark is so hard to detect or remove. At almost every point while generating text, a language model is choosing between several roughly equally good next words or phrases. Normally, that choice gets made using a random number. The watermarking method swaps that randomness for something that looks random but isn’t: a value derived from a secret key combined with the words generated so far.
The practical effect is that the model still produces high-quality, natural-sounding text, since it’s still picking from the same shortlist of good options it always would have. But the specific pattern of which options it picks, across a large enough stretch of text, forms a statistical fingerprint. A detector that knows the secret key can spot that fingerprint. A human reader never notices anything different, and the model itself isn’t even aware it’s happening.
Where the Watermark Falls Apart
There are real limits here:
- Heavy editing weakens or erases it. Since the watermark depends on the model’s specific word choices, substantially rewriting, paraphrasing, or translating the text can destroy the statistical pattern.
- Short text may not carry enough signal. A single sentence or two often isn’t long enough to reliably detect the pattern.
- Mixed-origin text is ambiguous. If you use an AI tool to proofread, translate, or lightly edit your own original writing, the output could still pick up a watermark, even though you wrote the substance of it yourself.
- Absence of a watermark proves nothing either way. Older AI models that predate this feature, or text that’s been heavily modified, may show no detectable watermark at all, without that meaning it wasn’t AI-assisted.
Quick Reference Table
| Company | Watermarking Approach | Status |
|---|---|---|
| Anthropic (Claude) | Invisible statistical watermark woven into word choice, plus C2PA metadata on files | Live for new models launched from August 2, 2026 |
| Google (Gemini) | SynthID, invisible watermarking across text, image, audio, video | Long-standing, applied across products |
| OpenAI (ChatGPT) | Transparency tools for images and audio via SynthID licensing | No public text watermark detection system confirmed yet |
| Detection tools | Third-party tools scan for statistical patterns, not just watermarks | Increasingly common, but imperfect |
Why This Was Rolled Out So Quietly
One detail stands out: none of this arrived with a big splashy announcement inside the actual chat interface you’re using. There was no in-app notice saying “by the way, everything you generate from now on carries a hidden signature.” It mostly surfaced through support pages, legal documentation, and reporting from tech journalists piecing details together.
Part of the push is regulatory. The EU AI Act’s Article 50 transparency requirements took effect August 2, 2026, and multiple major AI companies, including Anthropic, Google, and Meta, have signed onto the EU’s Code of Practice on Transparency of AI-Generated Content as a framework for compliance. But the rollout of these watermarks has generally extended worldwide, not just within the EU, suggesting the companies see business value in traceability beyond pure legal compliance.
How These Two Stories Connect
Here’s the bigger picture as a business owner. AI companies are dealing with a supply chain problem for clean, human-authored training data, and a transparency problem for their own generated output. Buying old books addresses the first. Watermarking generated text addresses the second, and arguably helps prevent AI-generated text from contaminating future training runs too, since properly tagged AI content can theoretically be filtered out of future scraped datasets.
For anyone publishing content online, this points to a future where the distinction between human-written and AI-generated content becomes increasingly detectable, whether you disclose it yourself or not. Treating your content strategy as though this distinction doesn’t matter is becoming a riskier bet by the month. If your site is built to be visible to AI search in the first place, that foundation also makes authentic content easier for both people and machines to verify.
What This Means Practically For Your Business
1. Assume Your AI-Assisted Content Is Traceable
Whether you’re using Claude, Gemini, or another major AI tool, assume that heavily unedited AI output could carry a detectable signal now or in the near future. This doesn’t mean stop using AI tools, it means stop publishing raw, unedited AI output as if it’s invisible.
2. Original, Human-Authored Content Just Became More Valuable
If AI companies themselves are paying real money to get their hands on guaranteed human-written text, that’s a strong signal about where value sits. Original writing, real customer stories, and first-hand expertise are becoming a more differentiated asset, not a less relevant one, and they’re also the raw material of a higher trust score for your site overall.
3. Editing Matters More Than Ever
Since heavy editing weakens or removes detectable watermarks, and since well-edited content also tends to be better content anyway, treating AI output as a first draft rather than a finished product serves two purposes at once: better quality, and less risk of publishing traceable, unedited AI text.
4. Disclosure Is Becoming the Safer Default
Rather than hoping nobody notices AI assistance in your content, consider being upfront about it where it matters, particularly for anything client-facing, YMYL topics like health or finance, or content where authenticity is part of your value proposition. Our E-E-A-T checklist covers the experience and trust signals that make disclosure easier to stand behind.
5. Diversify Your Content Creation Process
Mix AI-assisted drafting with genuine interviews, original data, and real subject matter expert input. This produces stronger content regardless of watermarking concerns, and builds a body of work that’s clearly rooted in real expertise.
What This Means for Local Businesses and Startups
If you’re running a local business or an early-stage startup relying on content marketing, here’s the grounded takeaway. You don’t need to panic about invisible watermarks buried in your blog post drafts. But you should treat this as more confirmation of something that was already true: content built on genuine local knowledge, real customer relationships, and authentic expertise holds up better than generic AI output, regardless of what’s happening behind the scenes with watermarking technology.
For local businesses specifically, this is a good moment to double down on original photography, real customer testimonials, and specific, local details in your website and Google Business Profile content, the exact kind of material that’s both hardest to fake and least likely to raise any transparency concerns down the line. We’ve written before about why your Google Business Profile is becoming a homepage for local search, and the authenticity argument only gets stronger from here.
How the F9XR Team Can Help
Navigating a landscape where AI companies are simultaneously hunting for clean training data and quietly tagging their own output takes more than guesswork. Business owners need a content strategy that holds up regardless of how AI detection technology evolves.
The F9XR Team helps business owners, startups, and local businesses build resilient content and digital strategies for exactly this moment, including:
- Content audits to identify thin or unedited AI text that could carry detectable signals or simply underperform with real audiences
- Editorial workflows that combine AI efficiency with genuine human expertise, original data, and authentic local knowledge
- Local SEO strategy built around real, original content that Google and AI search engines both reward
- Website development and website redesign work grounded in E-E-A-T principles from the ground up
- Ongoing digital presence management so your content strategy stays ahead of how AI platforms are evolving, rather than reacting after the fact
If you’re unsure whether your current content mix leans too heavily on unedited AI output, that’s a conversation worth having now.
Key Takeaways
- AI companies are buying pre-2022 printed books in bulk, sometimes over a million at a time, through brokers like ISBNdb, specifically because these books are guaranteed to be free of AI-generated text.
- The main drivers are avoiding model collapse, the documented quality decline that happens when AI models train on synthetic, AI-generated data, and avoiding intentional data poisoning tools authors have started using to fight back against unauthorized AI training.
- Scanning these books at industrial scale typically destroys them, which is why AI companies buying them often operate under strict NDAs and undisclosed identities.
- Separately, and connected to the same underlying problem, major AI companies including Anthropic and Google have quietly begun embedding invisible, statistical watermarks directly into AI-generated text.
- These watermarks work by subtly influencing which words a model chooses, forming a detectable pattern without changing the text’s quality or meaning, though heavy editing can weaken or erase the signal.
- Business owners should treat AI-generated content as a first draft requiring real human editing, and should recognize that original, human-authored content is becoming a more valuable, differentiated asset, not a less relevant one.
Conclusion
The fact that AI companies are simultaneously buying up old, printed books to avoid their own AI slop and quietly watermarking their own generated text tells you almost everything you need to know about where content is heading. Genuine, human-authored writing is becoming a scarcer, more valuable resource, even to the companies building AI itself, and AI-generated text is becoming more traceable, not less, as detection technology matures.
For business owners, the smartest response isn’t panic, it’s adaptation. Treat AI tools as a starting point, invest in genuine expertise and original content, and stop assuming AI-generated text is invisible or untraceable. If you want help building a content and digital strategy that holds up regardless of how AI detection technology evolves, teams like the F9XR Team work with business owners and local brands on exactly this kind of website development, website redesign, and local SEO strategy every day.
Produced using AI-assisted research and drafting workflows, then reviewed and edited by the F9XR editorial team. See our Editorial Policy for how we create and verify content.
Further Reading
Related Questions
Keywords
Copyright
© 2026 F9XR Team. Licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt this work with attribution to the F9XR Team.
Disclaimer
This article reflects the experience and opinions of the F9XR Team and is provided for general informational purposes only. It is not professional legal, financial, or medical advice. Metrics and results mentioned are illustrative and not guaranteed.
F9XR/Articles, published by the F9XR Team, delivers high-quality marketing content produced by staff and contracted experts, unless otherwise noted. This commitment to quality ensures comprehensive coverage of industry topics from a trusted team of writers. For more information, please visit the F9XR Team website or reach out through the contact page.
Ready to Build a Stronger Digital Presence?
The F9XR Team helps business owners turn websites into conversion engines. From website development and redesign to local SEO and AI-powered visibility, we engineer systems that bring in results, not just traffic.
Comments