, , ,

Anthropic challenges class certification in heavily redacted brief concealing information related to datasets used

On April 17, Anthropic filed its opposition to the plaintiffs’ motion for class certification. A large portion of Anthropic’s brief is redacted.

One intriguing argument by Anthropic is that the plaintiffs cannot identify a class due to the inability to ascertain whose books were used by Anthropic in their books datasets. The plaintiffs are proposing 2 classes: (1) Internet Books Class referred to by Plaintiffs as Pirated Books Class and a (2) redacted second class.

Anthropic argues: “The Court should also deny class certification because, as discussed in detail above, Plaintiffs present no viable method to ascertain who owns the copyrights to the books at issue. But to identify the rightsholders, one first needs to identify the books at issue, and Plaintiffs do not present a reliable or practical way to do this either: None of the methods proposed by Plaintiffs’ expert objectively or reliably identify the relevant books, much less their owners.

Without knowing which books datasets Anthropic is referring to, it’s hard to evaluate this argument. But one wonders why it’s so easy to identify the books Meta apparently used from books datasets it downloaded, such as LibGen. The Atlantic created a search tool for that.

At least according to the complaint, Anthropic allegedly used the Pile dataset, including Books3.

Leave a Reply


Discover more from Chat GPT Is Eating the World

Subscribe now to keep reading and get access to the full archive.

Continue reading