One response to “Video of oral argument in Doe 1 v. Github before 9th Circuit”
( Court properly in session at 6:04; https://youtu.be/50c-umo9uFU?t=364 )
I found the Microsoft defendant lawyer to have lost the plot slightly but he didn’t have time to get around to the more relevant material for the question at hand, choosing instead to make his first argument that plaintiffs don’t have standing. While that may be the case, that should be argued in the original court in order to have their claim dismissed should their argument survive the 9th.
The OpenAI defendant lawyer did make a concise good case with regard to having to potentially grapple with the notion that if *any* ‘copying’ violates 1202b without any of the requisites normally required for a copyright *infringement* case, it would open the floodgates for cases based on this alone.
Plaintiffs lawyer however also makes the perfectly reasonable case that there is little to no difference between copying a whole book minus the copyright disclaimer page, and copying every page but then deleting the copyright disclaimer page.
This is especially relevant with regard to the training stage, as defendant may have explicitly decided to exclude such pages as irrelevant / avoiding over-training, *or* have generally programmed their ingestion algorithm to avoid over-training based on patterns and the algorithm itself deciding that copyright disclaimer pages should be ignored. The reason it’s relevant is that in both scenarios, knowledge of the copyright disclaimer page occurred and whether it was removed-*by-removal* or removed-*by-omission* seems immaterial.
With regard to the output stage, I am not convinced that ‘removal’ can take place, even if an LLM or other model type regurgitates an entire work with staggering accuracy, based on my understanding of the technology.
However, it would be of interest for plaintiffs and/or researchers to ‘ask’ such a model to produce the next page after the end of the work (an author bio, perhaps) and if it produces that just fine, ask it to produce the page prior to the start of the work (a copyright disclaimer);
This would work in their advantage both ways:
If it *does not* produce the copyright disclaimer, it would strengthen the case that again there was a very deliberate ‘removal-by-removal’ of copyright disclaimers in general at the training stage, whether deliberate or otherwise.
But if it *does* produce the copyright disclaimer, the argument could be made that the model ‘knew’ or *could have* ‘known’ this when it regurgitated that entire work, but didn’t do so, strengthening a *removed-by-omission* argument.
Ultimately, the ‘solution’ arrived at by GitHub – which is to analyze an output and determine if it is substantially similar to another work, and attribute thusly – may be the only workable solution, at which point a ‘removal of CMI’ is no longer applicable.
Though I shudder to think such methods to be applied to *any* produced works, not just (code) LLM outputs. Analysis of a novella having attributions littered throughout its margins as parts of the written text are substantially similar to prior works, billboard hot 100 entries having lists of 50+ co-writers/composers as parts are found to be substantially similar, movie credits being an hour long (up from the 10 minutes to deal with the visual effects companies alone).
Where is the line drawn? Right now: “what line?”
Where *will* the line be drawn? I’m not sure the 9th has been given the tools in existing legislation to appropriate answer *that* question, even if they manage to resolve the CMI issue with regard to the identicality claims.
One response to “Video of oral argument in Doe 1 v. Github before 9th Circuit”
( Court properly in session at 6:04; https://youtu.be/50c-umo9uFU?t=364 )
I found the Microsoft defendant lawyer to have lost the plot slightly but he didn’t have time to get around to the more relevant material for the question at hand, choosing instead to make his first argument that plaintiffs don’t have standing. While that may be the case, that should be argued in the original court in order to have their claim dismissed should their argument survive the 9th.
The OpenAI defendant lawyer did make a concise good case with regard to having to potentially grapple with the notion that if *any* ‘copying’ violates 1202b without any of the requisites normally required for a copyright *infringement* case, it would open the floodgates for cases based on this alone.
Plaintiffs lawyer however also makes the perfectly reasonable case that there is little to no difference between copying a whole book minus the copyright disclaimer page, and copying every page but then deleting the copyright disclaimer page.
This is especially relevant with regard to the training stage, as defendant may have explicitly decided to exclude such pages as irrelevant / avoiding over-training, *or* have generally programmed their ingestion algorithm to avoid over-training based on patterns and the algorithm itself deciding that copyright disclaimer pages should be ignored. The reason it’s relevant is that in both scenarios, knowledge of the copyright disclaimer page occurred and whether it was removed-*by-removal* or removed-*by-omission* seems immaterial.
With regard to the output stage, I am not convinced that ‘removal’ can take place, even if an LLM or other model type regurgitates an entire work with staggering accuracy, based on my understanding of the technology.
However, it would be of interest for plaintiffs and/or researchers to ‘ask’ such a model to produce the next page after the end of the work (an author bio, perhaps) and if it produces that just fine, ask it to produce the page prior to the start of the work (a copyright disclaimer);
This would work in their advantage both ways:
If it *does not* produce the copyright disclaimer, it would strengthen the case that again there was a very deliberate ‘removal-by-removal’ of copyright disclaimers in general at the training stage, whether deliberate or otherwise.
But if it *does* produce the copyright disclaimer, the argument could be made that the model ‘knew’ or *could have* ‘known’ this when it regurgitated that entire work, but didn’t do so, strengthening a *removed-by-omission* argument.
Ultimately, the ‘solution’ arrived at by GitHub – which is to analyze an output and determine if it is substantially similar to another work, and attribute thusly – may be the only workable solution, at which point a ‘removal of CMI’ is no longer applicable.
Though I shudder to think such methods to be applied to *any* produced works, not just (code) LLM outputs. Analysis of a novella having attributions littered throughout its margins as parts of the written text are substantially similar to prior works, billboard hot 100 entries having lists of 50+ co-writers/composers as parts are found to be substantially similar, movie credits being an hour long (up from the 10 minutes to deal with the visual effects companies alone).
Where is the line drawn? Right now: “what line?”
Where *will* the line be drawn? I’m not sure the 9th has been given the tools in existing legislation to appropriate answer *that* question, even if they manage to resolve the CMI issue with regard to the identicality claims.