Microsoft exec says AI training might be the "largest labor theft in human history"

Daniel Sims

Posts: 2,568   +77
Staff
In brief: Aside from concerns about human extinction or a massive financial bubble, one of generative AI's biggest controversies is how tech giants train their models. While OpenAI and Microsoft claim that training AI on copyrighted material falls under fair use, internal comments from the companies could suggest otherwise.

Brent Hecht, Microsoft's director of applied science, predicted that people worldwide would view large language models as theft on an unprecedented scale, according to court documents released Thursday. The comments appear to counter Microsoft and OpenAI's public defense against an ongoing lawsuit brought by news outlets.

In the case's latest filing, Hecht said generative AI could be the largest labor theft in human history while describing how Microsoft, OpenAI, and other AI companies train their models on massive amounts of material from the open internet. The practice lies at the heart of the copyright lawsuit that The New York Times filed in late 2023.

Microsoft and OpenAI argue that training AI models on articles from the NYT and other publications falls under fair use, similar to a student reading books. The defendants also claim that generative AI output based on those articles sufficiently transforms the original material.

However, Thursday's filing contains internal comments from OpenAI that highlight AI's ability to reproduce news articles. For example, OpenAI executive Greg Brockman admitted that ChatGPT can predict and complete sentences from NYT articles when working with the outlet's material. Hecht also argued that winning the lawsuit might require the defendants to "make a complete mockery of the idea of fair use."

OpenAI has also previously admitted that training AI without accessing copyrighted material is impossible. Former Meta executive Nick Clegg echoed that view last year, claiming AI would die almost immediately if copyright law were enforced against it.

Meanwhile, AI companies have continued training their models on reams of user data, articles, and other material. Meta drew controversy for training its AI on employee behavior, Microsoft's GitHub Copilot trains on user data unless users opt out, and Twitch only started allowing streamers to opt out after training on their material for years. AI developers are also quietly buying, scanning, and destroying millions of books to gather more training material.

Generative AI's impact on the news business model is another pillar of the NYT's lawsuit. Google and other search engines now provide AI-generated summaries of information from articles in response to user queries, potentially discouraging readers from visiting news websites. In this week's court filing, an OpenAI executive involved in ChatGPT's development said publishers face an existential threat from AI.

Permalink to story:

 
Copyright law is antiquated and has needed changing for decades.

It was created long before AI was even a concept and really shouldn’t be relevant for most of these claims.

 
What AI does is akin to recording every TV show available free with ads and then selling boxed sets of those shows while claiming it's OK because they were "freely available" and definitely not stealing all the future ad revenue for that content for themselves.

And the only viable response for all content is to go behind a paywall, thus the end of the free (thanks to ads) internet.
 
Honestly I think it is fair use. But the real question is, should it be?

I think the argument that if a person can read a book and learn from it, so can AI - is a valid argument given current laws. The question we should be asking is: Do we need new laws specifically to address this?

There are a lot of strong arguments on both sides of that question. On some level it reminds me of the TNG episode "Measure of a Man". Imagine if that episode ended with them saying Data isn't aloud to learn from things the way a normal human is because of IP laws.
 
Honestly I think it is fair use. But the real question is, should it be?

I think the argument that if a person can read a book and learn from it, so can AI - is a valid argument given current laws. The question we should be asking is: Do we need new laws specifically to address this?

There are a lot of strong arguments on both sides of that question. On some level it reminds me of the TNG episode "Measure of a Man". Imagine if that episode ended with them saying Data isn't aloud to learn from things the way a normal human is because of IP laws.
It's the reselling others content that is the issue. If I read books and then write a new one on the same topic from my brain it's different than if I just copy and paste all the parts of other books into my "new" book (that's stealing). AI copy and pasting from a lot of other books doesn't make it different.
 
It's the reselling others content that is the issue. If I read books and then write a new one on the same topic from my brain it's different than if I just copy and paste all the parts of other books into my "new" book (that's stealing). AI copy and pasting from a lot of other books doesn't make it different.
A human could read an article and plagiarize it just like an AI could write something that's not plagiarism. If someone uses AI to put out something that is basically plagiarism, there's already laws for that.

I'm not saying I'm for or against anything here. It's an interesting issue. I'd probably lean towards some limits catered to AI but I don't really think this administration is capable of implementing effective rules so it's a tough one.
 
I'd probably lean towards some limits catered to AI but I don't really think this administration is capable of implementing effective rules so it's a tough one.

As a Brit, watching from t'other side of "the Pond", it's got nothing to do with whether your current administration is CAPABLE of implementing rules. It's obvious to everyone over here that they have no intention of doing so.....so far as the Orange Man is concerned, it's open the taps and 'full steam ahead'. And what was all that guff about a "strong President with a high IQ"???

I don't know who "the Donald" was referring to in that statement, but he definitely wasn't talking about himself, that's for sure! 🤣


Miq. :neutral:
 
I wonder if it's bots giving you those down votes. The recently re-introduced down votes would be an interesting experiment for dead internet theory
lol, there are certain posters who automatically disagree with whatever I say.. but hey, as long as I know I’m right, why should I care about them?
 
Back