AI research paper by Timnit Gebru, Li Fei-Fei, along with its dataset, is at center of copyright suit by EVOX Productions. Was it fair use or will Stanford University be liable for copyright infringement?

One copyright case that has not received enough attention is EVOX Productions v. Stanford University. EVOX contends that its stock car images were included in 2 datasets compiled by Stanford researchers for AI research: (1) ImageNet and (2) a smaller “Stanford cars dataset.” Both datasets were allegedly posted by Stanford on webpages to share with other researchers.

the imagenet dataset

As I’ve written before, Stanford University Professor Li Fei-Fei, who is not named in the lawsuit, is widely credited for advancing AI deep learning by overseeing the compilation of this important ImageNet dataset that was used by AI researchers in computer vision in an annual ImageNet competition.The AlexNet paper by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, submitted to this competition, is also widely considered one of the most important research papers in the history of AI research. It is one of the most cited AI research papers ever. And, yes, that’s the same Geoffrey Hinton who later was awarded the Nobel Prize for his advancements in AI research. Both the ImageNet dataset and the AlexNet paper were instrumental in showing how increasing the size of datasets can improve the capabilities of AI models, in a strategy known as scaling.

stanford cars dataset

In its most recent filing for the Joint Case Management Statement, Stanford explains the origin of the smaller car dataset below. The 2017 paper entitled Using Deep Learning and Google Street View to Estimate the Demographic Makeup of Neighborhoods Across the United States was co-authored by leading AI researchers, including the lead investigator Timnit Gebru and Fei-Fei Li. My highlighting added:

And here is a list of all AI researchers on the paper, which lists the lead researcher as Timnit Gebru:

Excerpt of title of paper

If what the Stanford researchers did is not fair use as EVOX Productions contends, then all universities in the United States that conducted or are conducting any AI training of models, or any use of datasets of unlicensed copyrighted works of any kind, will likely be at risk of liability. I explain why in my latest article on the “Fair Use and the Origin of AI Training.”

I think that would be the wrong result and terrible for the United States’s national interest in maintaining its position as the world leader in AI. But, until the court resolves this case (and other cases involving AI research uses of datasets), it remains an open legal question.

And, if Stanford loses, it’s possible all universities in the United States would lose as well for every unlicensed use of copyrighted materials in AI research. Such a drastic result could be disastrous for AI research and development in the United States. If academic research to advance the state of the art of one of the most transformative technologies of the 21st century is not the type of “research” that Section 107 was meant to foster, then it’s hard to imagine what is.

DOWNLOAD THE JOINT CASE MANAGEMENT STATEMENT:

Related Story:

Leave a Reply


Discover more from Chat GPT Is Eating the World

Subscribe now to keep reading and get access to the full archive.

Continue reading