Two more newsrooms are suing OpenAI. It's complicated.

Publishers are suing AI vendors, accusing them of theft. But are they giving the creators responsible for building their value a fair deal?

Link: Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft, by Anthony Ha in TechCrunch

I have some complicated feelings about these ongoing newsroom lawsuits against AI companies. The latest comes from The Seattle Times and Newsday. From the lawsuit itself:

“AI products like ChatGPT and CoPilot are touted as producers of content, but in fact they are rapacious consumers, devouring human-authored content and delivering back to the world copies and derivative imitations of that same original content they consumed to achieve their commercial objectives.”

The government has argued that training LLMs is fair use, and organizations like the EFF have agreed with it. The EFF’s argument in particular is important to take note of:

“The “market dilution” theory would eviscerate not only the fair use doctrine, but also other limits on copyright that work specifically to prevent rightsholders from unfairly suppressing competition by claiming broad ownership over tropes, genres, styles, and so on. In other words, publishers would wield unchecked veto power over any expression that might conceivably compete with a work they own.”

That’s a genuine problem! Expanding the scope of copyright law really does affect free expression. It’s an important concern to bring up — if the market dilution argument held, it could protect all sorts of businesses, both small and multinational.

No industry deserves to exist for its own sake, and we can’t ask the world to stay still to protect an industry. Systems and methods that work against journalism can’t be illegal in their own right. At the same time, we need journalism. An informed voting population is a prerequisite for a functioning democracy. The implication is that if the old methods no longer allow it to be sustainable, journalism must evolve and find new methods that are. We need information, context, and reporting; we don’t need the businesses that produce those things to operate in the exact same way they do now, or to look similar to how they do today.

There’s no doubt that AI in its current form takes work that is often independently produced by people with relatively few resources and uses it to enhance models that are controlled by large tech companies with billions of dollars at their disposal, which are often, in turn, controlled by billionaires. It’s a giant transfer of property from people with relatively low power to some of the most powerful people in the world. I think it’s fair to say they’re strip-mining culture to enrich themselves.

But those power dynamics are not entirely cut and dry. The EFF’s point is, in part, that expanding copyright scope will primarily benefit large corporations like Disney that have often, themselves, applied a chokehold to cultural expression. In preventing strip-mining by some billionaires, there’s a risk of giving new powers to other billionaires.

Although I want to focus on the power dynamics, there’s also a mechanical objection to consider. LLMs are performing statistical analysis on source material in order to train, and the outputs are created using probabilistic math, not intentional reproduction. We need to retain the ability to analyze creative work for all kinds of reasons, and even index it — consider search engine indices, which I think we generally want to exist so that people can find our work.

Of course, actual verbatim outputs — or approximate copies thereof — are plagiarism. And the materials LLMs are trained on are often obtained illegally or outside the terms of their licensing agreements: unambiguous theft that needs to be treated accordingly.

The Anthropic copyright settlement was about this: it wasn’t that training on books was found to be infringing in itself. In fact, Anthropic was found not liable for this. It was that it stole them. Anthropic settled for billions of dollars. That’s good in itself, but it turns out — surprise, surprise — that many publishers took the money for themselves and didn’t pass them down to authors, even when the rights had reverted to those authors.

The elephant in the room is that copyright was never designed to protect individual authors: it was designed to benefit the public, with author protections a necessary means to that end. In practice, most journalists, artists, etc, don’t own their work to begin with, and royalties are frequently corporation-to-corporation transfers even without AI.

The Newsday and Seattle Times lawsuits seek to protect corporate interests in a rapidly-changing market, with journalist reward effectively a trickle-down benefit. The licensing deals emerging as the market solution — among them OpenAI with News Corp, Axel Springer, the FT; Amazon with the New York Times — are corporate deals. Individual journalists and freelancers whose work constitutes the value see little of it, except in that it’s revenue that theoretically allows their employers to continue to exist and employ them. We probably need a new deal for creators that touches several levels.

The first is contracts and labor conditions. Unions have a huge part to play on this side of the equation. The Writers Guild of America won reasonable AI concessions from studios after their 2023 strike, ensuring that writer compensation couldn’t be diluted if a studio used AI. A revised contract this year also gave them written notice rights when studios intend to train models on their work, and a right to discuss compensation.

That’s great for members of the WGA, but most creative workers aren’t a part of an industry union. If that changed, both publishers and AI vendors would need to negotiate more creator-favorable terms. It seems like it should be a goal.

The second is infrastructure. We need independent creators to be able to set rights over their work and enforce them. Really Simple Licensing is one effort that’s working in this direction: a way to create an ASCAP-like standard for creative work. Notably, though, not a single major AI vendor has agreed to comply. There are criticisms of the collective licensing approach generally, and there’s still a platform compensation problem: if Reddit implements RSL, it will receive revenue, but its users who created the licensed content probably won’t.

The third is law. Our existing rules need to be updated to protect creators, regardless of AI; the advent of generative AI has made it even more important. Publishers must not be able to take all the money from licensing creative work (the EU’s Digital Single Market Directive is a good model to follow). AI vendors should be transparent about what they’re training on and how. Nobody should be able to steal training material. And rules should be revised to reflect that work can now be trained on and used to generate content at scale, automatically, which has knock-on effects on the public good.

These are all big topics. The shift we have to contend with is seismic in a way that hasn’t been rivaled since the advent of the web itself (and in some ways the impact on creators and publishers may be bigger than that). Each newsroom lawsuit is a kind of shot across the bow; the substantive conversations we need to be having are more expansive. But, in a political environment where AI vendors are aligned with the administration, and in a world where so many people want to be excited about AI in not very nuanced ways, I wonder if the conversations and actions that would lead to more equitable treatment and compensation for creators will be allowed to happen at all. Newsrooms should protect themselves, but they should make sure they’re looking out for the people who create their value, too.