People Matters Logo

Microsoft exec called AI scraping ‘the largest theft of labour in human history,’ unredacted filing reveals

• By Anjum Khan
Microsoft exec called AI scraping ‘the largest theft of labour in human history,’ unredacted filing reveals

A Microsoft executive described the use of scraped content to train AI models as “the largest theft of labour in human history”, according to newly unredacted material in The New York Times’ copyright lawsuit against Microsoft and OpenAI.

The comment, made by Microsoft Director of Applied Science Brent Hecht in a 2023 internal memo, is among previously redacted statements cited by the Times in its summary judgment filing. The documents also reveal concerns within Microsoft and OpenAI about AI’s potential impact on publishers, journalists and other content creators.

The lawsuit, filed in 2023, alleges that OpenAI and Microsoft used millions of copyrighted news articles to train AI systems without permission. The companies argue that their use of copyrighted material is protected by fair-use law. The newly unsealed material does not establish copyright infringement, and some underlying exhibits remain sealed.

Microsoft warned of an AI ‘doom loop’

A January 2024 Microsoft presentation warned that AI products could reduce traffic to publishers while weakening the content ecosystem that AI models rely on.

Hecht described this as a potential “doom loop”. Data cited in the filing showed that click-throughs from Microsoft’s Copilot to the Times’ website were as much as 93 per cent lower than from traditional Bing search.

Another Microsoft document warned of a “real risk” that generative AI could significantly disrupt the jobs of people whose work generated training data. Microsoft has said Hecht’s comments reflected his individual views rather than the company’s position.

OpenAI saw risks for publishers

Internal OpenAI communications also raised concerns about the impact of AI on news organisations. Nick Turley, who led the ChatGPT team, described publishers as facing an “existential threat” and said AI products were “largely substitutive”. 

OpenAI co-founder Greg Brockman said the models were particularly effective at news, while Microsoft CEO Satya Nadella acknowledged that AI could provide information directly rather than sending users to the original source.

The filings also allege that OpenAI researcher Nick Ryder told Brockman about a “hack” to bypass the Times’ paywall, to which Brockman responded, “ah nice”. Nadella separately testified that paywalled material should be licensed for AI training and said he would have required retraining had he known such content had been used.

Millions of news works allegedly included

The filing details the scale of content allegedly used in OpenAI’s training datasets. It says mid-training datasets contained more than 91,692 copies of works from the Times, Daily News and the Center for Investigative Reporting, while a Common Crawl-derived dataset contained more than two million documents from nytimes.com.

Project Mango allegedly contained at least 160,903 unique works from news publishers. The filing also describes alleged efforts to remove copyright notices from training data.

Copyright dispute remains unresolved

The case centres on whether using copyrighted works to train AI models without permission qualifies as fair use under US law.

Microsoft and OpenAI argue that AI training is transformative and does not substitute for the original works. The Times and other publishers argue that AI products can compete with journalism by providing information directly to users and reducing publisher traffic.

A recent US government brief backed OpenAI’s position on AI training, but that filing does not determine the outcome of the copyright dispute. The court has yet to rule on whether the companies’ training practices constitute infringement.