The latest filings in the New York Times matter against OpenAI and Microsoft don’t argue about text anymore. They invoke culture. They invoke sports. They invoke the full width of what human beings make when they think nobody is keeping score, which is a much larger category than a newspaper archive and a very different kind of claim. That shift in scope is the real story, and it is worth separating from the question of who eventually prevails. The scope tells you what the material was understood to be all along: not a licensed corpus, not a dataset anybody assembled on purpose, but the ambient output of a civilization that happened to be reachable.
Taking or Inheriting
The distinction that has to get counted, eventually and by somebody, is the difference between taking and inheriting. The internet worked because people loaded it with signal under the assumption that they were publishing, not that they were feeding a machine. Every review, every tutorial, every forum answer written at midnight for a stranger was contributed into what looked like a commons and turned out to be an input. That assumption is now the battlefield, and the reason it is worth writing about has nothing to do with who wins. It has to do with what the assumption was worth, and to whom.
Every Other Industry Pays
Anthropic is separately facing claims that it trained Claude on tens of thousands of songs without a license, and music is the cleanest place to see the shape of the problem, because music is already licensed. That infrastructure exists and has existed for a long time. The labels know how to clear a sample, the paperwork is routine, the rates are known, and the entire apparatus was built precisely so that new work could be made out of old work without anybody having to pretend the old work came from nowhere. The AI labs chose to treat that infrastructure as optional. That choice reveals everything about how these systems were built, because nobody skips a step that well established by accident. You skip it because you have decided in advance what the material is. The operating assumption was that the web was a commons, and that everything in it was raw material for the next model.
The scale of the taking is what makes this different from what came before, and scale is doing more work in that sentence than it appears to. Libraries pay for books. Radio stations pay for licenses. Sampling cleared its legal frame decades ago, and it cleared it the hard way, by arguing the point out in public until the terms were settled and the clearing became routine. The AI labs built their products fast enough and large enough that the accounting step felt unnecessary, and speed was the whole justification. By the time anyone thought to ask, the labs had already trained models, shipped products, and collected profits, which means the question arrives after the value has been realized rather than before. The question now is whether speed alone justifies skipping the step that every other industry treats as an ordinary cost of doing business.
The Deal Nobody Offered
In normal business, the first question for any strategic relationship is what the other side gets and what it would cost to replace them. You run the lifetime customer exercise. You study the replacement cost. You ask what investment the relationship demands going forward, because a supplier you intend to depend on for a decade is a different line item than a one-time purchase. None of that calculation ever happened for the training data, and the omission is strange precisely because these are companies that run the exercise for everything else. The writers, the photographers, the coders, and the musicians, the people who filled the web with signal in the first place, were never asked what their work was worth. The model simply took it. Nobody sent an invoice. Nobody negotiated terms. The opt-out button never mattered, and it never could have, because an opt-out is not a price. It is permission to be absent from a market you were never invited to participate in, offered after the fact by the party that set the terms.
Run the replacement-cost question honestly and the answer is uncomfortable for the labs. What would it cost to rebuild that corpus from scratch, commissioned and paid for at market rates? The number is large enough that nobody wants it on a slide, which is itself the admission. A supplier whose replacement cost is that high is not a commons. It is the most valuable relationship in the business, and it is the only one that was never formalized.
Copyright was supposed to be the mechanism that lets a person own the fruits of their own mind. That is the entire bargain: make something, hold it, decide who may use it and on what terms. Treat creative work instead as ambient data with no source, and you erode the conditions under which anybody bothers to make the distinctive thing in the first place, because the distinctive thing and the generic thing now earn the same return, which is none. A rule that degrades the person who made the work is unjust on its own terms, whatever it does for the balance sheet downstream. You cannot hold both positions at once: a society that rewards originality, and a training regime that assumes originality is free.
The Bill Comes Due
None of this ends with a simple win or a simple loss, and anybody promising you a clean outcome is selling something. What it does instead is force the question that got avoided during the training runs, when avoiding it was cheap and convenient: who owns the output when the input was everybody? That question does not have a technical answer. It has a commercial one, and the commercial answer is the one that will end up governing how these systems are built, whatever anyone decides on paper.
The companies that answer it honestly, meaning the ones that price the input and put the relationship on a footing the supplier would recognize as a relationship, will build models people trust and keep supplying. The ones that don’t will keep winning on benchmarks while losing the thing that makes the work worth doing, and they will lose it slowly enough to mistake the decline for weather. That is the part worth watching. Benchmarks measure the model. They do not measure whether the people who produce the next decade of signal still believe that publishing and feeding a machine are different acts.
The bill always comes due. The only question is whether you planned for it, or whether you are surprised to see it.

Leave a Reply