AI Frontier Post
AI News

USA Today sues OpenAI over AI training data: 160,000 articles in WebText, $250M sought

USA TODAY Co. sued OpenAI in federal court on October 8, 2026, alleging hundreds of thousands of its articles were scraped for AI training — including 160,000 entries in WebText — and seeking damages in excess of $250 million.

USA TODAY Co. — the company behind USA Today, the Detroit Free Press, the Indianapolis Star and 16 other newspapers — sued OpenAI in federal court on October 8, 2026, alleging the company trained its GPT models on hundreds of thousands of its copyrighted articles and now serves that journalism back inside chatbot answers that substitute for the originals. The suit seeks damages in excess of $250 million.

The filing: SDNY, today

The case, USA Today Co., Inc. v. OpenAI Foundation, was filed in the U.S. District Court for the Southern District of New York by attorney Steven Lieberman of Rothwell, Figg, Ernst & Manbeck, according to a docket-level account of the filing. The complaint was entered today alongside an exhibit of copyright registrations and a second exhibit of GPT-5.6 output examples; same-day docket entries include the civil cover sheet, a copyright notice form, a notice of appearance, and a request for issuance of summons. A statement of relatedness asks that the case be treated as related to the consolidated OpenAI copyright infringement litigation already pending in the same court.

The damages math follows the statutory framework the complaint cites: up to $150,000 for each willful infringement, plus up to $25,000 per violation for the removal of copyright management information. Named defendants include OpenAI Foundation and the for-profit entities — OpenAI GP, OpenAI OpCo, OpenAI Global, OAI International, OAI Corporation, and OpenAI Group PBC. The complaint also recounts OpenAI’s October 2025 recapitalization, under which the nonprofit became OpenAI Foundation holding equity valued at roughly $130 billion while the for-profit became OpenAI Group PBC; it cites OpenAI’s own disclosures for ChatGPT’s scale — more than 900 million weekly users and 50 million paying subscribers as of March 2026. The models at issue span GPT-1 through GPT-6.1 and GPT-OSS, including Instant, Thinking, mini, nano and Pro variants. The plaintiffs demand a jury trial.

The numbers: 160,000 articles in WebText

The complaint’s evidentiary centerpiece is counting. It says the plaintiffs’ content comprises more than 160,000 entries in WebText, the internal corpus OpenAI built to train GPT-2 — including 83,266 entries from usatoday.com and 12,994 from freep.com — and more than 122 million tokens in C4, the filtered English-language subset of a 2019 Common Crawl snapshot. It reproduces OpenAI’s published GPT-3 training mix, which weighted Common Crawl at 60 percent and WebText2 at 22 percent.

The training-data allegations go further than copying counts. The complaint alleges OpenAI scraped copyrighted material regardless of paywalls or access restrictions, and used programs designed to strip away the copyright management information marking the works as protected. It cites the written evidence OpenAI submitted to a British House of Lords inquiry in December 2023, in which the company said limiting training data to public-domain works would not produce AI systems that meet current needs. And it quotes internal communications to argue OpenAI knew what its models contained: co-founder Greg Brockman telling colleagues the models were particularly good at predicting news article text; a 2020 presentation by then-research leader Dario Amodei listing news generation among GPT-3’s skills; the VP of Research saying “We train our networks to memorize the training data — that’s their objective”; and June 2022 internal messages acknowledging that GPT-4 would have memorized large amounts of data and would be highly effective at regurgitating it.

The Thurgood Marshall United States Courthouse in Manhattan, home to the Southern District of New York.
The Thurgood Marshall U.S. Courthouse in Manhattan, home to the Southern District of New York, where the suit was filed. Photo: Pexels.

Project Taxi, Project Mango — and an “accidental cover-up”

Microsoft appears in the complaint in a supporting role with two codenames. The filing alleges that over a three-year period Microsoft provided OpenAI with a copy of the Bing Index — billions of scraped webpages including the plaintiffs’ content — under an initiative codenamed Project Taxi, and that Microsoft separately developed and operated a crawler called Project Mango on OpenAI’s behalf, for which OpenAI paid. It further alleges that OpenAI’s output filters did not suppress content from any entity that had not sued the company — an approach the complaint says a Microsoft executive described internally as an accidental cover-up.

The substitution theory: GPT-5.6 summaries

The second exhibit is the complaint’s answer to the fair-use question. It reproduces examples in which GPT-5.6, prompted to find and summarize a specific article by title, produced extensive multi-section summaries paraphrasing the originals and following their structural organization. The examples span papers across the company: the Indianapolis Star, the Detroit Free Press, The Knoxville News-Sentinel, The Palm Beach Post, The Tennessean, The Enquirer, The Des Moines Register, The Courier-Journal, the Naples Daily News, the Asbury Park Press, the Milwaukee Journal Sentinel, The Columbus Dispatch, The Oklahoman, and the Star News.

The complaint alleges OpenAI affirmatively post-trained its models to summarize copyrighted articles instead of returning the articles themselves — producing substitutes, it argues, that serve the same informative purpose as the originals. It quotes OpenAI’s Head of ChatGPT acknowledging that once ChatGPT gives an answer there is “no good reason to click” a link to the underlying source, and cites an OpenAI engineer’s statement that no matter how prominently links are displayed, users will not click them. The consequence, the plaintiffs allege: no reason to visit the original sources, and no reason to pay for a subscription.

A newspaper printing press running.
The plaintiffs hold copyrights across 19 publications, from USA Today to the Arizona Republic. Image: StockCake.

Why this one lands differently

The plaintiffs are the heavyweight of American local journalism: all owned by USA TODAY Co., Inc., formerly Gannett, which the complaint describes as a media company tracing its history to 1906, publisher of hundreds of daily publications, winner of dozens of Pulitzer Prizes, employing hundreds of journalists across thirteen states. The 19 publications at issue include USA Today, The Tennessean, the Indy Star, The Bergen Record, The Enquirer, the Asbury Park Press, the Democrat & Chronicle, The Knoxville News-Sentinel, the Naples Daily News, The Oklahoman, the Milwaukee Journal Sentinel, The Columbus Dispatch, The Arizona Republic, The Courier-Journal, The Des Moines Register, the Detroit Free Press, The Detroit News, The Palm Beach Post, and the Star News.

The timing matters too. The suit lands one day after Emmerich Newspapers and other local outlets sued Microsoft and OpenAI over the same training-data theory, and inside a fair-use fight that keeps escalating: the Justice Department weighed in for fair use in September, publishers fired back, and unsealed filings have surfaced internal admissions about scraping. There is also an irony the industry will notice: USA TODAY Co. has been on both sides of the AI bargaining table — Perplexity signed a licensing agreement with Gannett (now USA TODAY Co.), one of the publisher deals the AI companies point to as the legitimate path. The complaint’s answer is that a license for one product is not a license for everything. Everything here is an allegation, not a finding — but the docket is now real, and the fair-use question is getting bigger every month.