Largest Theft of Labor

The Largest Theft of Labor in Human History. Says the Defendant: Artificial Intelligence Trends

The plaintiffs called the defendants’ use of their content perhaps the “largest theft of labor in human history.” They were quoting the defendant when they said it.

The information came from internal emails and analysis in The New York Times’ lawsuit against OpenAI and Microsoft.

The redacted 92-page filing, unsealed last Thursday, quote executives talking about the “gazillions” of dollars at stake in commercializing AI models while privately acknowledging the “existential threat” those products pose to publishers’ underlying economics. Examples:

  • The Introduction to the filing references Microsoft’s director of Applied Science, Brent Hecht, who is quoted as referring to defendants’ actions as: “an astonishing theft of unprecedented proportions” and perhaps the “largest theft of labor in human history.”
  • OpenAI’s head of ChatGPT, Nick Turley, is quoted as writing: “[p]ublishers” face an “existential threat” from those products, which, “are largely substitutive, period” and “will get more and more substitutive as they get better.”
  • Around 2017, OpenAI co-founder Greg Brockman is quoted as having written that he was “deeply motivated by the gazillions” he hoped to gain by commercializing OpenAI’s technology.

Paywall Violations and CMI Removals vs. Fair Use Claims

One of the plaintiffs’ more compelling arguments was the claims that “Defendants used…web crawlers or scrapers…to access Plaintiffs’ webpages and extract their contents.” The filing noted: “OpenAI did not check whether Plaintiffs’ content was behind a paywall, or whether the content was subject to terms and conditions restricting use…For its part, Microsoft scraped Plaintiffs’ content ostensibly for its traditional Bing Search product, and then secretly transferred that content to OpenAI for model development and training” as part of “self-described ‘horse-trading’ deals” between the defendants.

The filing also references an instance “when [OpenAI researcher] Nick Ryder informed Mr. Brockman about ‘a hack to get around nytimes paywall’ to help with Mr. Brockman’s efforts to ‘scrape[]’ The Times’s site, Mr. Brockman responded ‘ah nice.’”

I’ve seen several articles covering the filing discuss the plaintiffs’ claims of paywall violations – and its impact on defendants’ “fair use” claims – but I haven’t seen them discuss plaintiffs’ claims that defendants removed copyright management information (“CMI”) applied under the Digital Millennium Copyright Act (DMCA). According to the filing, “OpenAI removed CMI from Plaintiffs’ Asserted Works without Plaintiffs’ permission…at a massive scale”. How massive? Here’s a table that plaintiffs included in their filing:

Source: News Plaintiffs’ Combined Summary Judgment Brief

This has led plaintiffs to state that: “The uncontested record establishes that the Court should enter summary judgment that statutory damages, if awarded, should be on a per-article basis.”

Pretty compelling stuff. Of course, we haven’t seen defendants’ request for summary judgment yet, though it has reportedly been filed – the judge will decide what other documents should be released to the public. But when a defendant’s representative is quoted as referring to defendants’ actions as perhaps the “largest theft of labor in human history”, the fair use argument could face an uphill battle.

So, what do you think? Does the plaintiffs’ filing impact your opinion as to whether defendants’ use of their content was “fair use”? Please share any comments you might have or if you’d like to know more about a particular topic.

Image created using ChatGPT, using the term “robot lawyer stepping around a wall that says ‘Paywall’ on it.”. Ironic, isn’t it? 😉

Disclaimer: The views represented herein are exclusively the views of the author, and do not necessarily represent the views held by my employer, my partners or my clients. eDiscovery Today is made available solely for educational purposes to provide general information about general eDiscovery principles and not to provide specific legal advice applicable to any particular circumstance. eDiscovery Today should not be used as a substitute for competent legal advice from a lawyer you have retained and who has agreed to represent you.


Discover more from eDiscovery Today by Doug Austin

Subscribe to get the latest posts sent to your email.

Leave a Reply