, ,

With statutory damages, the Shadow Library Strategy offers the same potential amount of recovery to book authors. Winning on AI training adds $0.

This post is a follow up to yesterday’s analysis of the Susman Godfrey Playbook in lawsuits against AI companies.

To recap, one of the strategies that Susman Godfrey lawyers successfully advanced in Bartz v. Anthropic is the separate treatment of Anthropic’s initial acquisition and library-building of copies from shadow libraries. What I call the Shadow Library Strategy.

Judge Alsup analyzed this initial acquisition and library-building as separate from Anthropic’s later use of some copies to train its AI model Claude. (Judge Chhabria disagreed with this approach in Kadrey v. Meta, and viewed the acquisition for the further purpose of training the AI model.)

On the surface, Judge Alsup’s decision looks like a split decision. He all but ruled that Anthropic’s (i) initial acquisition was copyright infringement, while Anthropic’s later (ii) training of its AI model was a fair use. The latter looks like a loss for the book authors.

Yet, in practical terms, the amount of statutory damages that the Bartz book authors would have been entitled to receive was exactly the same—meaning had they prevailed on both the initial acquisition and the training issue, if Judge Alsup had ruled it was all infringement, the Bartz book authors would have been entitled to only one recovery for each work infringed, no matter how many times infringed.

As the D.C. Circuit explained,

Both the text of the Copyright Act and its legislative history make clear that statutory damages are to be calculated according to the number of works infringed, not the number of infringements. 17 U.S.C. § 504(c) (1) authorizes a judge to award damages “for all infringements … with respect to any one work.”

As the House Report on the bill explains, however, only one penalty lies for multiple infringements of one work. See H.R.Rep. No. 1476, 94th Cong., 2d Sess. at 162 (1976), U.S.Code Cong. & Admin.News 1976, pp. 5659, 5778 (“A single infringer of a single work is liable for a single amount … no matter how many acts of infringement are involved in the action and regardless of whether the acts were separate, isolated or occurred in a related series…. Moreover, although the minimum and maximum amounts are to be multiplied where multiple ‘works’ are involved in the suit, the same is not true with respect to multiple copyrights … or multiple registrations.”) (emphasis added).

Walt Disney Co. v. Carl Powell, 897 F.2d 565, 569 (D.C. Cir. 1990)

Of course, the Bartz book authors didn’t have to elect statutory damages had the case gone to trial. Instead, they could have tried to prove actual damages, such as for (i) a lost license to train and (ii) potentially an apportioned defendant’s profits from the unauthorized use. But this probably would have been worse than pursuing statutory damages.

The lost license for the use of a book probably gets something close to what the proposed settlement amount is ($3,000 per book) based on the nascent market. So, on this issue, statutory damages probably would have been more appealing.

And, assuming that an AI company is making any profits right now (instead of burning through lots of cash), the potential for getting indirect profits, or what the defendant’s profits attributable to the plaintiffs’ works’ contribution, would likely present formidable challenges of proof. In the 9th Circuit, the plaintiff must show a “causal nexus between the infringement [of the work] and the gross revenue.Polar Bear Productions v. Timex Corp., 384 F.3d 700, 711 (9th Cir. 2004).

When a large language model has been trained on many millions of works, it probably will be hard for any expert to “proffer some evidence … [that] the infringement at least partially caused the profits that the infringer generated as a result of the infringement.” Id. (quoting Mackie, 269 F.3d at 911).

Yes, it’s true that books were viewed as higher quality data for training of LLMs. But, for any single book, it’s probably difficult to show a causal nexus. Although he received flak for his comments, Mark Zuckerberg was probably on target in his suggestion that authors “overestimate the value” of their work among the vast amount of data used to train AI models.

The bottom line is that the plaintiffs in Bartz v. Anthropic would likely have pursued statutory damages had the case gone to trial — and the total potential amount of statutory damages would not have changed even if they had won on the AI training issue.

(*Granted, the amount the jury picked could have changed based on the different acts of infringement. But it’s debatable how the jury would react to hearing evidence related to AI training — some might view it as an extenuating circumstance to warrant a lower amount of statutory damages, while others might view it as an aggravating circumstance making the initial acquisition worse. My guess is that more would see it as an extenuating circumstance to help explain why the company was even using a shadow library. Hearing evidence only related to a shadow library probably would sound bad.)

One response to “With statutory damages, the Shadow Library Strategy offers the same potential amount of recovery to book authors. Winning on AI training adds $0.”

  1. The “library building” conclusion is one I find puzzling. Whenever there are many items, a database is uses to organize. This is common and best practice. The practitioner needs to track what they have accessed, success rates, failures, retries, duplicates, and dozens more attributes. This is done using a database.

    You can’t have 7 million individual text files. That’s impractical and impossible to accurately track and utilize.

Leave a Reply


Discover more from Chat GPT Is Eating the World

Subscribe now to keep reading and get access to the full archive.

Continue reading