A new court filing reveals internal communications and testimony from OpenAI and Microsoft executives as part of a consolidated copyright infringement lawsuit.
Key facts
- •The lawsuit includes the New York Times, the Daily News group, Ziff Davis, the Center for Investigative Reporting, and The Intercept.
- •Microsoft CEO Satya Nadella testified that paywalled content should be licensed and stated he would have forced a model retrain had he known about the use of such data.
- •Internal Microsoft documents describe a "doom loop" where AI products threaten the economic foundations of their essential content suppliers.
- •OpenAI co-founder Greg Brockman acknowledged in an internal message that the company's models were "excellent at news" and could predict sentences from the New York Times.
- •The plaintiffs are seeking billions of dollars in damages and are asking the court to reject the fair use defense.
The New York Times and several other media organizations have filed a joint summary judgment brief in US District Court, alleging that OpenAI and Microsoft engaged in unauthorized use of their content to train AI models. The 92-page filing incorporates internal emails, Slack messages, and sworn testimony disclosed during discovery. The plaintiffs argue these documents undermine the defendants' fair use defense and demonstrate that AI products are actively substituting for original journalism.
By the numbers
Internal Concerns and Market Impact
The brief highlights internal warnings from Microsoft and OpenAI staff regarding the impact of AI on publishers. Microsoft director of applied science Brent Hecht described the practice of using such data as an "astonishing theft of unprecedented proportions." Similarly, OpenAI's head of ChatGPT, Nick Turley, characterized the products as "largely substitutive" and an "existential threat" to publishers. Microsoft's own data indicated that click-through rates on its Copilot tool were significantly lower than those on traditional Bing search, with declines ranging from 51% to 94% for the plaintiffs' publications.
Allegations of Paywall Bypassing
Plaintiffs allege that OpenAI systematically bypassed paywalls and ignored terms of service during its data collection process. The brief notes that OpenAI employees discussed "a hack to get around" the New York Times paywall, and the company's corporate representative testified that there was no method in place to detect or remove paywalled content from training data. Furthermore, the filing claims OpenAI used the "New York Times Annotated Corpus," which was licensed only for non-commercial research, despite internal acknowledgment that such use was inappropriate.
Fair Use and Licensing Arguments
The plaintiffs argue that the defendants' fair use defense is invalid because their AI models are commercial and substitutive rather than transformative. They point to the fact that OpenAI and Microsoft have signed licensing deals with other publishers, which the plaintiffs claim proves that a functioning licensing market exists and that the defendants chose to take content for free. Additionally, the brief notes that OpenAI shelved a tool called "Media Manager" in 2024 that was intended to allow publishers to opt out of scraping.
Timeline
- 2017Greg Brockman wrote about his motivation to commercialize OpenAI's technology.
- December 2023The New York Times filed its initial lawsuit against OpenAI and Microsoft.
- May 2025The US Copyright Office concluded that fair use cannot apply broadly to AI data scraping.
Advertisement
This article was independently rewritten by ManyPress editorial AI from reporting originally published by The Decoder.


