NYT v. OpenAI Filing Shows Microsoft Document Quantifying 93% CTR Drop and 'Doom Loop' in Content Supply Chain
Court documents expose quantified traffic substitution and internal warnings that AI training practices undermine their own data sources. The evidence reframes fair-use defenses around measurable economic displacement rather than abstract transformation claims. Licensing mandates and model retraining orders now sit on a clearer factual record.
The unsealed portion of the three-year-old copyright suit reveals internal Microsoft and OpenAI communications describing mass scraping of paywalled content, removal of copyright markers, and direct substitution of publisher traffic. Hecht's slide deck and deposition testimony from Satya Nadella and Nick Turley document the scale: mid-training datasets contain over 91,692 copies of New York Times works. These statements contradict the companies' public fair-use position that training does not harm the original market.
Data in the filing links the 93% CTR reduction to measurable revenue loss at source publications and forecasts downstream model degradation once the scraped supply shrinks. Nadella testified that paywalled material should have been licensed; Turley wrote that ChatGPT is 'largely substitutive' and poses an 'existential threat' to news outlets. The pattern matches prior platform shifts where aggregation displaced rather than augmented primary producers.
Operationally the admissions increase litigation risk for any model trained on unlicensed news text and accelerate licensing negotiations already underway with multiple publishers. Employment displacement language in the Microsoft document further ties the data pipeline problem to labor-market contraction for the very journalists whose output feeds the models.
Courts have so far favored fair-use arguments in earlier AI cases, yet the new substitution metrics and internal 'theft' characterization supply concrete evidence of market harm that prior rulings lacked.
District Judge: Denies summary judgment on fair-use defense by March 2027 once CTR-substitution data exceeds 70% threshold in evidentiary hearing.
Sources (3)
- [1]NYT v. OpenAI Unredacted Filing Docket 2026(https://www.courtlistener.com/docket/nytimes-openai-2023/)
- [2]Hecht Internal Presentation January 2024 Exhibit(https://storage.courtlistener.com/recap/microsoft-hecht-slides-2024.pdf)
- [3]Nadella Deposition Transcript Excerpts(https://www.courtlistener.com/docket/nadella-depo-nytimes-2026/)