, , ,

Tremblay plaintiffs seek discovery re: torrents from OpenAI

Magistrate Judge Illson denied several discovery motions of the Tremblay book author plaintiffs.

Yet, one interesting revelation from the order is that the plaintiffs are seeking discovery related to the word “torrents” from OpenAI.

Here’s the Judge’s discussion:

As to RFP No. 45 (Use of Torrents), Plaintiffs concede that OpenAI did not and is not excluding documents containing the term “torrent” while searching for “non-privileged documents
discussing the acquisition of text training data used to train the relevant models” and that it has produced some documents responsive to the request. Id. at 3. However, Plaintiffs contend that it “appears that other types of these documents have not been produced, such that a narrow, targeted request for this category of documents is warranted and proportional to the needs of the case.” Id. Plaintiffs then allude to “gaps” in Defendants’ production as a basis for a targeted ESI search. The court finds the bases for this request to be speculative and unsupported. Plaintiffs have failed to show why the current production is insufficient.

In an unrelated case, Kadrey v. Meta, Meta’s apparent use of torrents that included seeding of some files for others to download has become a major issue.

Also interesting is what Magistrate Judge Illson ordered OpenAI to do:


(1) its search for complaints about copyright issues, its documents concerning efforts to prevent regurgitation of training materials, and its productions concerning agreements and negotiations with third parties will not exclude documents because they concern a GPT- class model in development;

(2) it will produce for inspection the text pre-training data for in-development, text-based GPT-class of models where the pre-training phase has been completed and the model is intended for production, including the next GPT-class model still being developed, which has been referred to as “Orion”; and

(3) should Plaintiffs have questions about the origin of a particular dataset, or whether that set was legally licensed, they may address them to OpenAI.

Leave a Reply


Discover more from Chat GPT Is Eating the World

Subscribe now to keep reading and get access to the full archive.

Continue reading