When I spoke at the NeurIPS conference a couple years ago, I prodded the audience full of AI researchers from universities to think about their roles in the ongoing debate over the acquisition and use of copyrighted materials to research, develop, and train AI models.
I didn’t say it there but said it in my law review article Fair Use and the Origin of AI Training: “The practice—what some call, invoking religious terms, AI’s ‘original sin’—began in university research and later migrated to companies after the AI research showed promise.”
In other words, if being trained on unlicensed works is AI’s original sin, AI researchers at universities committed it.
Based on my historical research, Jack Bandy and Nicholas Vincent’s 2021 scholarly article questioning the “documentation debt” in the compilation and use of the BookCorpus based on copyrighted books may be the first serious discussion among university AI researchers about whether what AI researchers were doing was a violation of copyright.
They wrote: “Are there tasks for which BookCorpus should not be used? We leave this question to be more thoroughly addressed in future work. However, our work strongly suggests that researchers should use BookCorpus with caution for any task, namely due to potential copyright violations, duplicate books, and sampling skews.”
I don’t know how much academic discussion their article provoked at universities. But my very limited sense is probably not enough.
What about fair use?
Of course, one possibility is that what university AI researchers did and are doing is a fair use.
Even without much knowledge of fair use doctrine, university AI researchers may have assumed that “scholarship” and “research” — which are expressly mentioned in the fair use provision — cover their own AI research within fair use.
Yet, that quite reasonable position on fair use will soon be tested in at least some of the more than 100 copyright lawsuits against AI companies and one even against Stanford University for its university AI researchers’ compilation of the influential ImageNet dataset on copyrighted images. Many of the plaintiffs argue categorically that acquisition and use of copyrighted materials to train is not fair use or transformative.
Which cases to follow?
These are the 4 copyright cases to watch closely:
These 4 cases all involve research or scholarship in ways that implicate the larger question whether university AI researchers will need to license any copyrighted materials used in AI training and research.
This includes AI researchers using any AI models trained by others on unlicensed materials, especially if they involve memorization of any of the training materials. Even researchers researching AI memorization may not be immune from potential lawsuits that implicate the practices of their universities. Oy vey.
The Non-Zero Chance University AI Researchers Are at Risk
I wish we weren’t in this state of affairs and already had court decisions clearly recognizing some fair use for this important research at universities. But we don’t.
The lawsuit against Stanford University — for what Stanford researchers did in compiling one of the most important datasets, ImageNet, in the history of AI research — even survived a motion to dismiss.
So here we are. The law is unsettled. Two district court judges ruled that AI training of LLMs by Anthropic and Meta were highly transformative fair uses. But even their opinions were unfavorable to the AI companies in other respects, including, in the Anthropic case, the provenance and retention of the data used.
I say the chance of copyright exposure for academic researchers is “non-zero,” which suggests it is small. But things can change quickly if, for example, the case went against Stanford University or one of the AI companies raising fair use based on research purposes (without a commercially deployed model).
“And the entire edifice on which all university-based AI research in the United States was built— datasets—may come crashing down like a house of cards.”
Related Stories