The federal system answered Aaron Swartz's 70 GB JSTOR download with criminal charges; Meta's 81.7 TB book trove for AI training is a civil fair use fight.
Federal prosecutors once brought a 35-year sentencing threat against a researcher for downloading 70 gigabytes of academic papers from JSTOR. Twelve years later, the same federal system is handling a consolidated civil action against Meta over an alleged 81.7 terabytes of books used to train its Llama family of AI models. The relevant copyright statutes are unchanged.
In early 2025, plaintiffs in Kadrey v. Meta (N.D. Cal., 3:23-cv-03417) unsealed filings alleging that Meta torrented at least 81.7 terabytes of books from shadow libraries to train its Llama family of AI models (Ars Technica). That figure is about 1,167 times the roughly 70 gigabytes federal prosecutors say Aaron Swartz downloaded from JSTOR in 2011.
Aaron Swartz, who in his twenties began systematically retrieving academic articles from JSTOR through MIT's network, faced federal charges under wire-fraud and computer-misuse statutes. The case carried a 35-year sentencing threat, a $1 million fine, and asset forfeiture (community discussion of the case record). JSTOR did not pursue civil litigation; the prosecution was federal. Swartz died by suicide in January 2013, two years into the case and before trial.
The Meta case is structurally different in law and in tone. Kadrey v. Meta, brought by authors including Sarah Silverman, Ta-Nehisi Coates, Richard Kadrey, and Christopher Golden, is a consolidated civil copyright class action (CourtListener docket). Meta's defense is fair use, and the matter proceeds as commercial litigation, not as a criminal referral. Newly unredacted filings in 2025 indicate that Mark Zuckerberg personally approved the Llama team's use of the LibGen-style dataset for training (TechCrunch). Vanity Fair reporting describes Meta AI staff categorizing more than 7 million books as having "no economic value" in internal systems (Vanity Fair).
The relevant statutes are not new. The Copyright Act, the Computer Fraud and Abuse Act, and the wire-fraud statutes the government used against Swartz were all on the books when Meta's team routed traffic through the same peer-to-peer protocol long associated with piracy. What is new is the scale at which they were used and the corporate authorization under which they were used. The plaintiffs' filings allege Meta's downloads spanned multiple shadow libraries, including at least 81.7 TB from Anna's Archive, 35.7 TB from Z-Library and LibGen, and 80.6 TB from LibGen on a prior occasion. The same filing argues that piracy cases involving a fraction of one percent of Meta's alleged volume have historically drawn far harsher consequences.
In October 2025, Meta filed a response in the same docket arguing that other bulk downloads on its own infrastructure IPs, including porn traffic, were for "personal use" rather than AI training (Ars Technica). The filing did not address the authors' training claims. The document makes a narrower point the firm is now testing in court: that bulk downloads by Meta personnel, on Meta infrastructure, for purposes Meta later defines, are categorically different from the same act performed by an individual researcher. The court has not yet ruled on that distinction.
A blog post that circulated earlier in the year framed the comparison as prosecutorial murder (curiousquail). The legal record does not support that framing. It does support a narrower claim: federal criminal copyright enforcement has not, in twelve years, scaled with the volume of infringement. It has scaled with the actor. Swartz faced federal charges for a download that, by Meta's own alleged figures, was less than 0.1% of the size of the dataset his former employer's peers have been accused of torrenting. The cases are not directly comparable, and Meta's fair-use defense has not been adjudicated. The cleanest name for the gap is selective reach, and the question the Kadrey docket is being asked to answer is whether that selectivity survives contact with a court.