
The Justice Department has taken sides in the escalating legal conflict between publishers and the artificial intelligence industry, arguing that developers such as OpenAI should be allowed to train their models on copyrighted news articles without paying licensing fees. In a court filing submitted on Tuesday in the Southern District of New York, the department said that treating such training as copyright infringement would not only slow the progress of American AI but also damage national security by advancing the position of foreign competitors.
The filing is a statement of interest rather than a direct ruling, but it carries significant weight. It signals that the US government is concerned about the broader geopolitical consequences of allowing copyright law to become an obstacle to AI development. The Justice Department argued that rules making it harder to build a robust American AI industry threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered.
The government's position aligns with OpenAI's fair-use defence in the consolidated publisher cases before the Southern District of New York. OpenAI has long maintained that training machine-learning models on publicly available text, including news articles, is a lawful transformative use under US copyright law. The company has repeatedly insisted that it does not reproduce protected expression in its outputs, and that the models learn patterns, facts, and language structures from the vast corpora on which they are trained.
The Justice Department went one step further, arguing that imposing licensing costs on AI model developers would risk handing the largest technology companies an oligopoly on model training. The reasoning is that only the wealthiest firms would be able to afford the massive licensing fees required to train on copyrighted content at scale. Smaller startups and open-source developers would be priced out of the market, leaving a handful of dominant players in control of the next generation of AI infrastructure.
That argument has resonated with parts of the AI research community, which fears that broad copyright liability would centralise power in organisations with vast legal and financial resources. Supporters of the government's position say that a strongly protective copyright regime would discourage experimentation and slow the United States' ability to maintain leadership in a technology widely viewed as strategically important.
Publishers, however, reject the framing. The New York Times, one of the leading plaintiffs in the consolidated litigation, responded sharply to the Justice Department's filing. Spokesperson Graham James said the administration is siding with a handful of trillion-dollar AI companies and that AI companies need to pay fairly for the content that makes their products possible. For news organisations, the issue is existential: they invest heavily in journalism, only to see AI systems trained on their reporting and, in some alleged instances, generate textual content that closely resembles their original articles.
The New York Times sued OpenAI and its primary financial backer, Microsoft, in December 2023, accusing the companies of using millions of copyrighted articles without permission. Ziff Davis, the media company that owns CNET, filed its own lawsuit in 2025. More than 400 local newspapers are pursuing separate action. The cases have been consolidated before the same federal judge, making the outcome particularly important for the future relationship between the AI industry and the news business.
The publishers say that OpenAI and other developers are building profit-making products on the back of journalism produced at great expense. They argue that fair use was never intended to allow wholesale commercial exploitation of copyrighted works by technology companies. They also point to instances in which ChatGPT and other large language models have reproduced, nearly verbatim, extended passages from paywalled articles, undermining both the value of subscription models and the ability of publishers to control their own content.
OpenAI counters that its systems are trained on data that includes a broad cross-section of the internet and that the resulting models generate entirely new text. The company has also made deals with numerous publishers in an effort to defuse the conflict, paying some outlets for access to their content and for permission to display their works in AI-generated responses. Those licensing agreements have not resolved the underlying legal question, and several publishers have chosen to litigate rather than negotiate.
But even if the publishers lose in the United States, the legal landscape is very different on the other side of the Atlantic. European copyright law has no fair-use doctrine. Instead, it relies on a closed list of exceptions, meaning that uses of protected works without authorisation are permitted only in narrowly defined circumstances. The absence of a general fair-use principle is not accidental; it reflects a broader European approach that places greater emphasis on protecting the rights of authors and publishers.
What Europe has instead is a text-and-data-mining exception with an opt-out mechanism. Under Article 4(3) of the 2019 Directive on Copyright in the Digital Single Market, rightsholders can expressly reserve their rights over the use of their works for text and data mining. If they do, the exception stops applying to those works. This opt-out system has become a crucial battleground for AI developers operating in Europe, because without a clear licence or a failure to reserve rights, they may not lawfully mine copyrighted content.
The European Union's Artificial Intelligence Act plugs directly into that mechanism. Article 53 requires every provider of general-purpose AI models to adopt a policy that identifies and respects reservations of rights made by rightsholders under the copyright directive. In practice, this means that an AI developer must keep up to date with opt-out declarations and ensure that any training data obtained from published sources is filtered accordingly.
Importantly, the obligation is attached to the model itself, not merely to the original training run. Recital 106 of the AI Act states that any provider placing a general-purpose AI model on the Union market must comply with the copyright obligations, regardless of the jurisdiction in which the copyright-relevant acts took place. The recital also says that no provider should be able to gain a competitive advantage in the Union market by applying lower copyright standards than those prevailing in Europe.
That proviso turns the Justice Department's national-security argument on its head. The Justice Department wants US law to shield AI developers from broad copyright liability in order to maintain American competitiveness. The EU, by contrast, wants to prevent AI companies from exploiting weaker copyright regimes outside Europe to undermine the rights of European creators. In effect, Brussels is insisting that any AI model sold or deployed in the EU must respect the copyright opt-outs recorded by European publishers, no matter where the training data was collected.
The consequences of this transatlantic divergence are already visible in the court rulings coming out of Europe. In November, a Munich court ruled against OpenAI in a copyright case involving song lyrics. The court found that lyrics memorised in the GPT-4 model amounted to reproduction of protected works and that the text-and-data-mining exception did not cover the use. The judgment was particularly notable because it rejected the argument that AI training on copyrighted text is permissible merely because it happens as part of an automated computational process.
That ruling is under appeal and could eventually reach the Court of Justice of the European Union, which would provide authoritative guidance on how EU copyright law applies to large language models. If the Munich judgment is upheld, AI developers will face even tighter constraints in Europe, especially if they are unable to demonstrate that they have respected reservation-of-rights declarations from rightsholders. Publishers in Europe have already accused OpenAI of withholding evidence about its training data and of failing to comply with its transparency obligations, adding another layer of complexity to the proceedings.
The result is a fast-growing legal patchwork. In the United States, a decision in favour of OpenAI could remove one of the biggest financial risks facing the AI industry, clearing the way for companies to continue training models on vast quantities of copyrighted material without payment. In Europe, even a clear American victory would not grant OpenAI immunity, because EU law applies to any provider placing a model on the Union market. The same model that is legal in California could remain unlawful in Berlin or Paris.
This means that a win for OpenAI in Manhattan would not travel. The company would still face its European obligations in a market where publishers have shown increasing willingness to use the courts and the regulatory system to demand accountability. The US Justice Department's national-security argument may resonate in Washington, but it has little force in European capitals, where copyright protection is treated as a fundamental prerequisite for a healthy creative economy.
For AI developers, the practical takeaway is that the legal risks vary sharply by region. A single global model trained without regard for copyright reservations may be tolerated under US fair-use principles but could expose the provider to liability in the EU. For publishers, the stakes are equally high: they are not only fighting over past training data but also over the rules that will determine how AI companies use news content in the future. With cases moving through courts on both sides of the Atlantic, the relationship between journalism and artificial intelligence is being redefined by judges, rather than by the tech companies or the publishers alone.
