A legal argument now getting wide attention holds that obtaining copyrighted books through BitTorrent, and even participating in their distribution, can qualify as fair use when the end purpose is training a large language model. The public framing of that argument is binary and systemic: either bulk scraping is legal and data stays close to free, or it is not and the cost of building frontier models jumps.
That framing bundles three separable questions into one headline, and it assumes the outcome most of the industry is rooting for is the outcome that leaves the industry safest.
Three separable questions
A training data claim breaks into parts that fail independently:
- Training. Does using a lawfully held copy to fit model weights count as transformative? This is the argument with the most doctrinal support and the most academic sympathy.
- Acquisition. How did the copy get onto the disk? Whether a transformative downstream use launders the method of obtaining the input remains unsettled.
- Distribution. Did the acquiring party also serve copies to other people? On BitTorrent this is not a hypothetical. Most clients upload pieces to peers while downloading by default. Seeding ratios and behavior vary by client, configuration, and network conditions, but downloading a torrent is often also an act of sharing it.
Distribution is doctrinally the weakest of the three. Redistributing exact copies of books wholesale to strangers resists a transformative use defense because it substitutes directly for the original and engages one of the core exclusive rights copyright law protects. A company can win on training, lose on distribution, and still owe a very large number.
That is the first thing the binary framing hides. There is a third branch, where courts bless the training theory and still penalize the acquisition and seeding conduct, and it carries its own liability profile.
Winning is the expensive outcome
For most companies in this ecosystem, a clean permissive ruling is the worse long-term result.
A restrictive precedent is painful and legible. It forces a licensing market into existence, and licensing costs are a line item. Line items get negotiated, insured, and disclosed in diligence.
A permissive ruling on training leaves the acquisition and distribution exposure intact while suppressing market attention to it, so the risk stays off the balance sheet as the underlying corpus grows. Statutory damages in United States copyright generally scale with the number of works infringed, though per-work awards remain discretionary within statutory bands and are subject to aggregation arguments for compilations and joint works. Depending on registration timing and whether a court treats a corpus as a single compilation, exposure attached to a large corpus scales with dataset size rather than model revenue. The precise arithmetic depends on jurisdiction and findings about willfulness, so treat that as a shape rather than a forecast.
Correlation makes it worse. A handful of publicly documented book and image corpora are widely inferred to sit underneath a large share of released models. That creates correlated, though not identical, exposure across defendants with different acquisition methods and jurisdictions. An unpriced, correlated, non-diversifiable exposure sitting under an entire capital stack is systemic fragility, because it stays invisible to auditors until an adverse ruling lands.
A permissive ruling also validates completed acts of scraping, which likely favors incumbents. It hands a late entrant a doctrine rather than a corpus, since many of the most cited sources have since been taken down, paywalled, or contractually fenced. Free data as a legal doctrine and free data as an available resource are different things.
The mitigation is clerical
If acquisition and distribution can fail while training succeeds, the useful hedge is bookkeeping. What separates a bounded exposure from an unbounded one is whether a company can reconstruct, per document, where a training example came from and under what terms. In practice:
- Keep acquisition logs that survive team turnover, including the retrieval method, the date, and the source endpoint.
- Treat method of acquisition as a separate field from license status. A public URL is not a license, and a permissive license on a wrapper does not cover the contents.
- Never use a protocol that uploads by default for anything you are not entitled to distribute. That is a configuration setting someone can fix today.
- Segment corpora so that a bad finding about one source does not require retraining everything.
- Price the licensing scenario now, even at a wide range, so that a restrictive ruling becomes a budget event instead of an existential one.
Jurisdictions will diverge. Some regimes have text and data mining exceptions with opt-out mechanics that do not map onto United States fair use analysis at all, so a single ruling will not settle the global picture even if it is decisive at home.
Here is the claim I will defend. The industry is treating a multi-branch legal question as a coin flip, and it is treating its preferred branch as the safe one. That branch keeps a large, correlated liability off the books for another few years. It defers the bill rather than cancelling it.