Court documents cited in reporting indicate that executives at OpenAI and Microsoft discussed the risks of using publisher material to train ChatGPT. The material came from web scraping and included millions of news articles, as first reported by CNET.
The records reportedly capture concerns at both companies over that practice. One Microsoft executive described OpenAI’s use of publishers’ work as “the largest theft of labor in human history.” The comment puts unusually direct language on a dispute that has pitted publishers against AI developers over copyrighted training material.
Publisher training data faces growing legal pressure
Web scraping lets developers collect pages at a volume that would be difficult to assemble manually. The reported cache matters because it included millions of articles, putting newsroom output inside the large datasets used by the chatbot. Publishers have challenged the use of their copyrighted content for AI training, and the documents could inform arguments about whether developers can use reporting without permission or licenses.
Those cases could set terms for licensing arrangements between newsrooms and model makers. They also test whether public availability on the web permits copying at training scale. The reported internal concerns may matter to those disputes, especially as courts weigh how AI companies obtained and used the material behind their systems.
Microsoft’s role extends the stakes beyond ChatGPT
Microsoft is a major partner and investor in OpenAI, so its reported discussion carries weight beyond the creator of ChatGPT. The companies’ relationship ties the data question to a business alliance that has helped place the chatbot at the center of the generative AI market. For news organizations, the outcome of copyright challenges could influence how they control and seek payment for their archives.
The available summaries do not identify the executives, the court case, filing date or underlying documents. They also do not say which publishers or articles appeared in the scraped material, or whether the concerns prompted changes to policies or training data. Those gaps limit what can be concluded about the conversations and their effect on ChatGPT’s development.
This article was produced with AI assistance from multi-source reporting and is published under our editorial standards.