AI firms face backlash over unpaid training data
The article argues that large AI companies rely on books, news, code, and other online material to train models while resisting payment or license compliance. It cites unsealed court documents in the New York Times case against OpenAI and Microsoft, including internal Microsoft comments about data use and employment disruption. The piece contends this practice threatens the content ecosystem that supplies training material.
Unsealed filings in the New York Times case, covered by 404 Media, cite Microsoft scientist Brent Hecht calling the alleged data use an extraordinary theft and possibly history’s largest labor theft. A Microsoft policy document warned generative AI could displace the people whose data trained it, undermining its own supply chain and risking model collapse.
OpenAI and Microsoft argue fair use, saying public articles and books advance knowledge without replacing their market. A cited deposition says OpenAI made no known effort to detect or remove paywalled content; cofounder Greg Brockman reportedly welcomed bypassing a firewall. Judge Sidney H. Stein has not ruled.
If courts accept that training on copyrighted material without payment is fair use, writers, journalists, coders, and other creators may see fewer licensing opportunities and weaker bargaining power. Readers could face a narrower, more homogenized information landscape if independent outlets and authors struggle to fund original work. Conversely, clearer rules or paid licensing could support creators while shaping how AI systems are built. The case’s outcome may influence investment, access, and trust in AI tools.