US Government Backs OpenAI in Legal Fight Over Training LLMs with Copyrighted Works

The U.S. government has filed a 20-page brief supporting OpenAI in a lawsuit brought by The New York Times over the alleged unauthorized use of copyrighted material to train its large language models (LLMs). The filing argues that allowing continued use of copyrighted works is essential for the United States to preserve its lead in the rapidly evolving AI industry. The brief leans heavily on an executive order signed by President Donald Trump last year, which declared maintaining global leadership in artificial intelligence a national priority.

The case centers around whether training LLMs—with vast datasets including copyrighted books, articles, and other media—without explicit permission violates fair use doctrine. Publishers, including The New York Times, contend that OpenAI and similar companies should not be allowed to rely on such data without compensating creators. The government’s brief pushes back, saying that overly restrictive interpretations of fair use could disrupt scientific progress, economic growth, and the overall competitiveness of the U.S. tech sector.

Legal precedents in recent years have delivered mixed outcomes. For instance, a judge ordered Anthropic to pay $1.5 billion after writers sued over use of their works—but the ruling wasn’t based on the training process itself. Instead, Anthropic was penalized for accessing pirated shadow libraries to obtain materials.

While the government’s brief is not binding—because the Southern District of New York, where the case is underway, won’t be directly controlled by the executive—it still carries potential weight. Amicus filings like this can influence judicial thinking even if they aren’t part of the case’s jurisdiction.

Why It Matters

This dispute isn’t just about one lawsuit—it’s a flashpoint for how copyright law and AI advancement will intersect going forward. Fair use has generally been interpreted flexibly in cases involving AI training data, though courts are increasingly scrutinizing where the line should be drawn—especially when the data was sourced without authorization. Publishers argue for tighter limits or compensation; AI developers warn that restrictions could make it harder for models to reach the scale and performance needed for breakthroughs.

The government’s involvement signals its view that the economic stakes are high. With AI-originated industries poised for massive growth, any legal framework that curtails training methods could put U.S. companies at a disadvantage on the global stage. On the flip side, among stakeholders invested in creative content—authors, publishers, artists—concerns over compensation, attribution, and control of original works will continue to drive pressure for new legal clarity.

What to watch next: how the Southern District of New York rules on fair use in this case, whether Congress steps in with clearer copyright law updates, and how other countries respond. The outcome could set an early standard for what’s acceptable—legally, ethically, and economically—as AI models increasingly depend on vast stores of existing works.