The Wall Street Journal’s Joanna Stern interviewed CTO Mira Murati on Stern’s podcast.
There was a telling exchange in which Murati appeared unsure or evasive in answering what data OpenAI used to train Sora, its text-to-video generator, including being “unsure” about whether OpenAI is using YouTube videos.
Well, when asked about it, YouTube CEO Neal Mohan said he’s unsure about what was done, but, if OpenAI did train Sora on YouTube videos, that would violate YouTube’s terms of service.
Fascinating development. Of course, Google itself is being sued for copyright infringement, in JL v. Alphabet, for its training of its AI models on content allegedly scraped from the Internet. The Complaint alleged: “As part of its theft of personal data, Google illegally accessed restricted, subscriptionbased websites to take the content of millions without permission and infringed at least 200 million materials explicitly protected by copyright, including previously stolen property from websites known for pirated collections of books and other creative works. Without this mass theft of private and copyrighted information belonging to real people, communicated to unique communities for specific purposes, and targeting specific audiences, many of Google’s AI products including Bard would not exist. Defendant continues to feed its AI products stolen data through regular updates with new personal and protected information scraped from internet users without any consent.”
So, it would be complicated if YouTube, owned by Alphabet, contemplated a lawsuit against OpenAI if it is discovered Sora is being trained on YouTube videos.