The Wall Street Journal’s Joanna Stern interviewed CTO Mira Murati on Stern’s podcast.
There was a telling exchange in which Murati appeared unsure or evasive in answering what data OpenAI used to train Sora, its text-to-video generator.
Joanna Stern: What data was used to train Sora?
Mira Murati: We used publicly available data and licensed data.
Joanna Stern: So videos on YouTube?
Mira Murati: I’m actually not sure about that.
Joanna Stern: Okay. Videos from Facebook? Instagram?
Mira Murati: If they were publicly available, publicly available to use, there might be the data, but I’m not sure. I’m not confident about it.
Joanna Stern: What about Shutterstock? I know you guys have a deal with them.
Mira Murati: I’m just not going to go into the details of the data that was used, but it was publicly available or licensed data.
Joanna Stern: After the interview, Murati confirmed that the licensed data does include content from Shutterstock. Right now, Sora is going through red teaming, AKA, the process where people test the tool to make sure it’s safe, secure, and reliable. The goal is to identify vulnerabilities, biases, and other harmful issues. What are things that just you won’t be able to generate with this?